Skip to content

Deploy the multimodal service

Same shape as every recipe. No new IAM. Vision uses the same models + endpoints as E2 chat — the deploy story here is: compose E1 baseline + E2 call-an-llm and add nothing else.

Official docs verified 2026-08-08

  • Azure OpenAI vision uses the same Cognitive account as chat: learn.microsoft.com/…/openai/how-to/gpt-with-vision
  • Vertex Gemini vision uses roles/aiplatform.user — same role as chat.
  • Bedrock Converse image input uses bedrock:InvokeModel — the same policy the E2 module already emits.

Layout

examples/multimodal/
  service/                        # composes E2 chat_native + chat_claude
    analyze.py                    # per-cloud image-in-chat call
    main.py                       # FastAPI /analyze + /analyze-claude
    Dockerfile                    # layers call-an-llm/service alongside multimodal/service
    requirements.txt
  chart/                          # Helm chart matching the E2 env-var block
    Chart.yaml
    values.yaml
    templates/{deployment,service,serviceaccount}.yaml
  terraform/
    azure/  (module "baseline" + "call_an_llm" — nothing else)
    gcp/    (module "baseline" + "call_an_llm" — nothing else)
    aws/    (module "baseline" + "call_an_llm" — nothing else)

Terraform — pure composition

module "baseline" {
  source  = "../../../foundations/azure"
  name    = var.name
  region  = var.region
  compute = var.compute
}

module "call_an_llm" {
  source                    = "../../../call-an-llm/terraform/azure"
  name                      = var.name
  region                    = var.region
  compute                   = var.compute
  azure_openai_base_url     = var.azure_openai_base_url
  azure_openai_deployment   = var.azure_openai_vision_deployment   # e.g. "gpt-5"
  foundry_claude_deployment = var.foundry_claude_deployment
}

Vision runs on the same Azure OpenAI account + deployment as chat — just pick a vision-enabled model name for azure_openai_vision_deployment (gpt-5 is the current default; gpt-4.1, gpt-4o, and o-series reasoning models also accept image input).

module "baseline" {
  source     = "../../../foundations/gcp"
  project_id = var.project_id
  name       = var.name
  region     = var.region
  compute    = var.compute
}

module "call_an_llm" {
  source     = "../../../call-an-llm/terraform/gcp"
  project_id = var.project_id
  name       = var.name
  region     = var.region
  compute    = var.compute
}

Vertex Gemini vision uses the same roles/aiplatform.user binding the call-an-llm module already emits. No extra role, no extra service.

module "baseline" {
  source  = "../../../foundations/aws"
  name    = var.name
  region  = var.region
  compute = var.compute
}

module "call_an_llm" {
  source                = "../../../call-an-llm/terraform/aws"
  name                  = var.name
  region                = var.region
  compute               = var.compute
  eks_oidc_provider_arn = var.eks_oidc_provider_arn
}

Nova Pro / Nova Lite + Claude-on-Bedrock all invoke via bedrock:InvokeModel, which the E2 module's IAM policy already grants (scoped to arn:aws:bedrock:*::foundation-model/*).

Helm — unchanged from E2

The E2 chat chart's env-var block already includes every LLM key vision needs. The P2.4 chart adds one env — CHIRON_STACK (native or claude) — and a matching vision-model env for each provider:

helm upgrade --install multimodal ./examples/multimodal/chart \
  --set image.repository="…/multimodal" --set image.tag="v1" \
  --set env.CHIRON_PROVIDER="azure" \
  --set env.CHIRON_STACK="native" \
  --set env.AZURE_OPENAI_VISION_DEPLOYMENT="gpt-5"

Verify (validate-only)

for p in azure gcp aws; do
  ( cd examples/multimodal/terraform/$p && terraform init -backend=false && terraform validate )
done

helm lint examples/multimodal/chart
helm template multimodal examples/multimodal/chart > /dev/null

python -m compileall examples/multimodal/service

Live vision-chat needs real credentials + model access — Phase 3 covers billed apply.