Skip to content

Multimodal (vision input)

Image + text → LLM. Same chat SDK you already learned in Call an LLM — just a different content shape on the message body. The E2 chat service is the LLM here; multimodal only teaches how to feed it an image.

The pipeline

     ┌───────────────┐    ┌────────────────────┐    ┌──────────┐
     │  image bytes  │──► │  chat.create with  │──► │  text    │
     │   + prompt    │    │  image content     │    │  reply   │
     └───────────────┘    │  block (per SDK)   │    └──────────┘
                          └────────────────────┘
                            (same E2 SDK path,
                             different content shape)

Per-cloud shape

Every provider has one method + one content-block schema; the code below is the entire delta from a plain text chat call.

Cloud Model (default) Method Image content block
Azure gpt-5 (current default) / gpt-4.1 / gpt-4o / o-series client.chat.completions.create {"type": "image_url", "image_url": {"url": "…"}} (URL or data:image/…;base64,…)
GCP gemini-3.6-flash (current default) / gemini-3.5-flash / gemini-2.5-pro / gemini-2.5-flash client.models.generate_content types.Part.from_bytes(data=bytes, mime_type="image/jpeg")
AWS Amazon Nova (Pro / Lite / Micro) OR Anthropic Claude on Bedrock client.converse {"image": {"format": "jpeg" \| "png", "source": {"bytes": <raw bytes>}}} (no base64 wrap when using boto3)

Claude through-line

AnthropicBedrock / AnthropicVertex / AnthropicFoundry all take the same {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "<base64>"}} content block. One shape, three per-cloud clients — same pattern as the E2 chat through-line, only the content list is different.

Docs verified 2026-08-08

Vision-enabled model IDs (verified)

Provider Model IDs that accept image input
Azure OpenAI (per vision-enabled docs) gpt-5 series (current default), gpt-4.1 series, gpt-4.5, gpt-4o, gpt-4o-mini, o-series reasoning models
Vertex Gemini gemini-3.6-flash (current default), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-2.5-pro, gemini-2.5-flash — all GA per Gemini catalog 2026-08-08.
Bedrock native (Converse) amazon.nova-pro-v1:0, amazon.nova-lite-v1:0 (Nova Micro is text-only)
Bedrock Anthropic (Converse) global.anthropic.claude-opus-4-6-v1, us.anthropic.claude-sonnet-4-5-20250929-v1:0 — image via Anthropic SDK image block or Converse image block
Vertex Anthropic claude-opus-5, claude-sonnet-5, claude-haiku-4-5@20251001 (base64 sources only on Bedrock + Vertex per Anthropic docs)
Foundry Anthropic claude-opus-5, claude-sonnet-5, claude-haiku-4-5

Pages

  • Analyze an image — the request per cloud, native flagship + Claude through-line.
  • Deploy — Terraform composing E1 baseline + E2 call-an-llm; the Helm chart reuses the E2 env-var block unchanged (no new keys needed).

Validate-only

No live LLM calls in CI. Each page's "verify" block is the local one-liner.