Skip to content

Deploy the chat service

Package the chat code from Chat as a container and deploy it two ways per provider — without duplicating the E1 Terraform + Helm baseline.

Every Terraform module here uses module "baseline" { source = "../../../foundations/<provider>" } and layers only the LLM-specific bits on top (model-access IAM, service env vars). The Helm chart extends the E1 chart with the chat container's env-var set.

Official docs verified 2026-08-08

Layout

examples/call-an-llm/
  service/                        # Python chat service (native + Claude paths)
    chat_native.py                # per-provider native flagship
    chat_claude.py                # Anthropic through-line
    tools_loop.py                 # multi-turn tool-use example
    Dockerfile                    # single image, per-provider env vars pick the SDK
    requirements.txt
  chart/                          # Helm chart extending the E1 baseline
    Chart.yaml
    values.yaml
    templates/
      deployment.yaml
      service.yaml
  terraform/
    azure/     (module "baseline" { source = "../../../foundations/azure" ... })
    gcp/       (module "baseline" { source = "../../../foundations/gcp" ... })
    aws/       (module "baseline" { source = "../../../foundations/aws" ... })

Terraform — E1 baseline REUSE, not duplication

Each provider's module composes the E1 baseline:

module "baseline" {
  source  = "../../../foundations/azure"
  name    = var.name
  region  = var.region
  compute = var.compute      # "kubernetes" or "serverless"
}

On the serverless path, an azurerm_container_app extends module.baseline.container_apps_environment_id and injects AZURE_OPENAI_* env vars for the chat container. On kubernetes, the Helm install (below) receives the same env vars via values overrides.

module "baseline" {
  source     = "../../../foundations/gcp"
  project_id = var.project_id
  name       = var.name
  region     = var.region
  compute    = var.compute
}

On serverless, a google_cloud_run_v2_service in this module overrides the baseline's placeholder image with the chat container's real Artifact-Registry URL, and mounts GOOGLE_CLOUD_PROJECT + GEMINI_MODEL env vars. On kubernetes, Helm install (below) takes over.

module "baseline" {
  source  = "../../../foundations/aws"
  name    = var.name
  region  = var.region
  compute = var.compute
}

On serverless, an aws_apprunner_service reuses the baseline's aws_iam_role.apprunner_access (via the module's apprunner_access_role_arn output) and adds a task role with bedrock:InvokeModel on the target model ARN. On kubernetes, EKS + IRSA follows the same Helm-install flow as the other two clouds.

Helm — extending the E1 chart

The E1 baseline chart is provider-agnostic (deployment + service). The E2 chart at examples/call-an-llm/chart/ adds the chat container's env-var block and IRSA/workload-identity service-account annotations.

Install (once the cluster + registry are up):

helm upgrade --install chat ./examples/call-an-llm/chart \
  --set image.repository="$(terraform -chdir=examples/call-an-llm/terraform/azure output -raw container_registry_login_server)/chat" \
  --set image.tag="v1" \
  --set env.AZURE_OPENAI_BASE_URL="https://<res>.openai.azure.com/openai/v1/" \
  --set env.AZURE_OPENAI_DEPLOYMENT="gpt-5"
helm upgrade --install chat ./examples/call-an-llm/chart \
  --set image.repository="$(terraform -chdir=examples/call-an-llm/terraform/gcp output -raw artifact_registry_url)/chat" \
  --set image.tag="v1" \
  --set env.GOOGLE_CLOUD_PROJECT="$GCP_PROJECT" \
  --set env.GEMINI_MODEL="gemini-3.6-flash"
helm upgrade --install chat ./examples/call-an-llm/chart \
  --set image.repository="$(terraform -chdir=examples/call-an-llm/terraform/aws output -raw ecr_repository_url)" \
  --set image.tag="v1" \
  --set env.BEDROCK_MODEL_ID="amazon.nova-pro-v1:0" \
  --set env.AWS_REGION="us-east-1" \
  --set serviceAccount.annotations."eks\.amazonaws\.com/role-arn"="$(terraform -chdir=examples/call-an-llm/terraform/aws output -raw chat_role_arn)"

Serverless path

Serverless is a single terraform apply from the same directory with var.compute = "serverless". The relevant resources:

azurerm_container_app on the baseline's azurerm_container_app_environment — the chat container image from ACR, env vars mounted, scale-to-zero enabled.

google_cloud_run_v2_service — image from Artifact Registry, env vars in the containers.env block, min_instance_count = 0 for scale-to-zero.

aws_apprunner_service — image from ECR, runtime_environment_variables for env vars, source_configuration.auto_deployments_enabled = false (recipes deploy explicitly, not on ECR push).

Verify (validate-only)

# Terraform
cd examples/call-an-llm/terraform/<provider>
terraform init -backend=false
terraform validate

# Helm
helm lint examples/call-an-llm/chart
helm template chat examples/call-an-llm/chart > /dev/null

# Python
python -m compileall examples/call-an-llm/service

Live apply (Phase 3, deferred)

Same command inverted for teardown:

terraform destroy   # both k8s and serverless

Live LLM calls need real accounts + model access; Chiron E2 is validate-only. Phase 3 (billed apply against real clouds) lands in a later batch.