Multimodal (vision input)¶
Image + text → LLM. Same chat SDK you already learned in Call an LLM — just a different content shape on the message body. The E2 chat service is the LLM here; multimodal only teaches how to feed it an image.
The pipeline¶
┌───────────────┐ ┌────────────────────┐ ┌──────────┐
│ image bytes │──► │ chat.create with │──► │ text │
│ + prompt │ │ image content │ │ reply │
└───────────────┘ │ block (per SDK) │ └──────────┘
└────────────────────┘
(same E2 SDK path,
different content shape)
Per-cloud shape¶
Every provider has one method + one content-block schema; the code below is the entire delta from a plain text chat call.
| Cloud | Model (default) | Method | Image content block |
|---|---|---|---|
| Azure | gpt-5 (current default) / gpt-4.1 / gpt-4o / o-series |
client.chat.completions.create |
{"type": "image_url", "image_url": {"url": "…"}} (URL or data:image/…;base64,…) |
| GCP | gemini-3.6-flash (current default) / gemini-3.5-flash / gemini-2.5-pro / gemini-2.5-flash |
client.models.generate_content |
types.Part.from_bytes(data=bytes, mime_type="image/jpeg") |
| AWS | Amazon Nova (Pro / Lite / Micro) OR Anthropic Claude on Bedrock | client.converse |
{"image": {"format": "jpeg" \| "png", "source": {"bytes": <raw bytes>}}} (no base64 wrap when using boto3) |
Claude through-line¶
AnthropicBedrock / AnthropicVertex / AnthropicFoundry all take the same
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "<base64>"}} content block. One shape, three per-cloud clients — same pattern as the E2 chat through-line, only the content list is different.
Docs verified 2026-08-08¶
- Azure GPT vision — learn.microsoft.com/…/openai/how-to/gpt-with-vision
- Vertex Gemini image understanding — ai.google.dev/gemini-api/docs/image-understanding
- google-genai
types.Part.from_bytes— googleapis.github.io/python-genai - Bedrock Converse ImageBlock — docs.aws.amazon.com/bedrock/…/conversation-inference
- Anthropic vision — platform.claude.com/docs/en/build-with-claude/vision
Vision-enabled model IDs (verified)¶
| Provider | Model IDs that accept image input |
|---|---|
| Azure OpenAI (per vision-enabled docs) | gpt-5 series (current default), gpt-4.1 series, gpt-4.5, gpt-4o, gpt-4o-mini, o-series reasoning models |
| Vertex Gemini | gemini-3.6-flash (current default), gemini-3.5-flash, gemini-3.5-flash-lite, gemini-2.5-pro, gemini-2.5-flash — all GA per Gemini catalog 2026-08-08. |
| Bedrock native (Converse) | amazon.nova-pro-v1:0, amazon.nova-lite-v1:0 (Nova Micro is text-only) |
| Bedrock Anthropic (Converse) | global.anthropic.claude-opus-4-6-v1, us.anthropic.claude-sonnet-4-5-20250929-v1:0 — image via Anthropic SDK image block or Converse image block |
| Vertex Anthropic | claude-opus-5, claude-sonnet-5, claude-haiku-4-5@20251001 (base64 sources only on Bedrock + Vertex per Anthropic docs) |
| Foundry Anthropic | claude-opus-5, claude-sonnet-5, claude-haiku-4-5 |
Pages¶
- Analyze an image — the request per cloud, native flagship + Claude through-line.
- Deploy — Terraform composing E1 baseline + E2 call-an-llm; the Helm chart reuses the E2 env-var block unchanged (no new keys needed).
Validate-only¶
No live LLM calls in CI. Each page's "verify" block is the local one-liner.