Deploy the multimodal service¶
Same shape as every recipe. No new IAM. Vision uses the same models + endpoints as E2 chat — the deploy story here is: compose E1 baseline + E2 call-an-llm and add nothing else.
Official docs verified 2026-08-08
- Azure OpenAI vision uses the same Cognitive account as chat: learn.microsoft.com/…/openai/how-to/gpt-with-vision
- Vertex Gemini vision uses
roles/aiplatform.user— same role as chat. - Bedrock Converse image input uses
bedrock:InvokeModel— the same policy the E2 module already emits.
Layout¶
examples/multimodal/
service/ # composes E2 chat_native + chat_claude
analyze.py # per-cloud image-in-chat call
main.py # FastAPI /analyze + /analyze-claude
Dockerfile # layers call-an-llm/service alongside multimodal/service
requirements.txt
chart/ # Helm chart matching the E2 env-var block
Chart.yaml
values.yaml
templates/{deployment,service,serviceaccount}.yaml
terraform/
azure/ (module "baseline" + "call_an_llm" — nothing else)
gcp/ (module "baseline" + "call_an_llm" — nothing else)
aws/ (module "baseline" + "call_an_llm" — nothing else)
Terraform — pure composition¶
module "baseline" {
source = "../../../foundations/azure"
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/azure"
name = var.name
region = var.region
compute = var.compute
azure_openai_base_url = var.azure_openai_base_url
azure_openai_deployment = var.azure_openai_vision_deployment # e.g. "gpt-5"
foundry_claude_deployment = var.foundry_claude_deployment
}
Vision runs on the same Azure OpenAI account + deployment as chat — just pick a vision-enabled model name for azure_openai_vision_deployment (gpt-5 is the current default; gpt-4.1, gpt-4o, and o-series reasoning models also accept image input).
module "baseline" {
source = "../../../foundations/gcp"
project_id = var.project_id
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/gcp"
project_id = var.project_id
name = var.name
region = var.region
compute = var.compute
}
Vertex Gemini vision uses the same roles/aiplatform.user binding the call-an-llm module already emits. No extra role, no extra service.
module "baseline" {
source = "../../../foundations/aws"
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/aws"
name = var.name
region = var.region
compute = var.compute
eks_oidc_provider_arn = var.eks_oidc_provider_arn
}
Nova Pro / Nova Lite + Claude-on-Bedrock all invoke via bedrock:InvokeModel, which the E2 module's IAM policy already grants (scoped to arn:aws:bedrock:*::foundation-model/*).
Helm — unchanged from E2¶
The E2 chat chart's env-var block already includes every LLM key vision needs. The P2.4 chart adds one env — CHIRON_STACK (native or claude) — and a matching vision-model env for each provider:
helm upgrade --install multimodal ./examples/multimodal/chart \
--set image.repository="…/multimodal" --set image.tag="v1" \
--set env.CHIRON_PROVIDER="azure" \
--set env.CHIRON_STACK="native" \
--set env.AZURE_OPENAI_VISION_DEPLOYMENT="gpt-5"
Verify (validate-only)¶
for p in azure gcp aws; do
( cd examples/multimodal/terraform/$p && terraform init -backend=false && terraform validate )
done
helm lint examples/multimodal/chart
helm template multimodal examples/multimodal/chart > /dev/null
python -m compileall examples/multimodal/service
Live vision-chat needs real credentials + model access — Phase 3 covers billed apply.