Analyze an image¶
Send an image + text prompt to a vision-enabled model; get a text answer. The E2 chat call gets one extra content-part in the user message.
Official docs verified 2026-08-08
- Azure GPT vision — learn.microsoft.com/…/openai/how-to/gpt-with-vision
- Vertex Gemini image understanding — ai.google.dev/gemini-api/docs/image-understanding
- Bedrock Converse image block — docs.aws.amazon.com/bedrock/…/conversation-inference
- Anthropic vision — platform.claude.com/docs/en/build-with-claude/vision
Native flagship¶
import base64, os
from openai import AzureOpenAI
client = AzureOpenAI(
api_key=os.environ["AZURE_OPENAI_API_KEY"],
azure_endpoint=os.environ["AZURE_OPENAI_ENDPOINT"],
api_version=os.environ.get("AZURE_OPENAI_API_VERSION", "2025-04-01-preview"),
)
with open("input.jpg", "rb") as f:
b64 = base64.b64encode(f.read()).decode()
resp = client.chat.completions.create(
model=os.environ.get("AZURE_OPENAI_VISION_DEPLOYMENT", "gpt-5"),
messages=[
{"role": "user", "content": [
{"type": "text", "text": "Describe this picture:"},
{"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{b64}"}},
]},
],
max_tokens=2000,
)
print(resp.choices[0].message.content)
A publicly reachable image_url also works — swap the base64 data URI for the raw URL string.
types.Part.from_bytes is the load-bearing helper; the rest of the call is the same generate_content you already use for text-only chat.
import os
from google import genai
from google.genai import types
client = genai.Client(
vertexai=True,
project=os.environ["GOOGLE_CLOUD_PROJECT"],
location=os.environ.get("GOOGLE_CLOUD_LOCATION", "us-central1"),
)
with open("input.jpg", "rb") as f:
image_bytes = f.read()
resp = client.models.generate_content(
model=os.environ.get("GEMINI_VISION_MODEL", "gemini-3.6-flash"),
contents=[
types.Part.from_bytes(data=image_bytes, mime_type="image/jpeg"),
"Describe this picture:",
],
)
print(resp.text)
For a GCS-hosted image, swap Part.from_bytes for
types.Part.from_uri(file_uri="gs://…/image.jpg", mime_type="image/jpeg").
Converse takes the raw bytes directly — boto3 handles base64 encoding for you (per the Bedrock docs: "If you use an AWS SDK, you don't need to encode the bytes in base64").
import os, boto3
client = boto3.client("bedrock-runtime",
region_name=os.environ.get("AWS_REGION", "us-east-1"))
with open("input.jpg", "rb") as f:
image_bytes = f.read()
resp = client.converse(
modelId=os.environ.get("BEDROCK_VISION_MODEL", "amazon.nova-pro-v1:0"),
messages=[{
"role": "user",
"content": [
{"image": {"format": "jpeg", "source": {"bytes": image_bytes}}},
{"text": "Describe this picture:"},
],
}],
)
print(resp["output"]["message"]["content"][0]["text"])
Claude through-line (same shape on all 3 clouds)¶
Anthropic's image content block is one dict passed to messages.create(content=[...]) on any of AnthropicBedrock / AnthropicVertex / AnthropicFoundry. On Bedrock + Vertex the source type must be "base64" (URL sources aren't available there yet, per the Anthropic vision docs).
import base64
with open("input.jpg", "rb") as f:
b64 = base64.standard_b64encode(f.read()).decode()
msg = client.messages.create( # any of Foundry / Vertex / Bedrock
model=…,
max_tokens=1024,
messages=[
{"role": "user", "content": [
{"type": "image",
"source": {"type": "base64",
"media_type": "image/jpeg",
"data": b64}},
{"type": "text", "text": "Describe this picture:"},
]},
],
)
print(msg.content[0].text)
Anthropic guidance is that images work best before text in the content list, so the service snippets above lead with the image.
Common signature¶
The service exposes:
CHIRON_PROVIDER picks the backend; a CHIRON_STACK env var selects native flagship vs Claude through-line (defaults to native).
Multi-image¶
Every provider supports multiple images in one turn — just add more image content blocks. Anthropic's docs recommend labeling each one (Image 1:, Image 2:) so follow-up prompts can refer to them by name.