Embed¶
Turn text into a vector. Every RAG pipeline starts here; every provider has a per-cloud SDK path that produces a comparable-shape response.
Official docs verified 2026-08-08
- Azure OpenAI embeddings: learn.microsoft.com/…/openai/how-to/embeddings
- Vertex Gemini embeddings: ai.google.dev/gemini-api/docs/embeddings
- Bedrock Titan embeddings: docs.aws.amazon.com/bedrock/…/titan-embedding-models
Model IDs¶
| Provider | Model ID (default) | Dimensions | Notes |
|---|---|---|---|
| Azure OpenAI | text-embedding-3-small |
1536 | -large variant: 3072. Uses your DEPLOYMENT name. |
| Vertex AI | gemini-embedding-001 |
3072 (default) | Older stable IDs: text-embedding-005, text-multilingual-embedding-002. |
| Bedrock | amazon.titan-embed-text-v2:0 |
1024 (default), 512, 256 | Older stable ID: amazon.titan-embed-text-v1 (1536 dims). |
The dimension choice matters — the store schema has to match. All examples below hold to the default row of the table.
Native flagship¶
Uses the same openai package + openai/v1/ base URL as Call an LLM → Chat. model= is the DEPLOYMENT name of your embedding model.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AZURE_OPENAI_API_KEY"],
base_url=os.environ["AZURE_OPENAI_BASE_URL"], # …/openai/v1/
)
resp = client.embeddings.create(
model=os.environ.get("AZURE_EMBED_DEPLOYMENT", "text-embedding-3-small"),
input="the quick brown fox jumped over the lazy dog",
)
vector = resp.data[0].embedding # list[float], len 1536
import os
from google import genai
client = genai.Client(
vertexai=True,
project=os.environ["GOOGLE_CLOUD_PROJECT"],
location=os.environ.get("GOOGLE_CLOUD_LOCATION", "us-central1"),
)
resp = client.models.embed_content(
model=os.environ.get("GEMINI_EMBED_MODEL", "gemini-embedding-001"),
contents="the quick brown fox jumped over the lazy dog",
)
vector = resp.embeddings[0].values # list[float]
Bedrock embeddings are invoke_model (not Converse — Converse doesn't cover embeddings, per the Titan embed docs).
import json, os, boto3
client = boto3.client(
"bedrock-runtime",
region_name=os.environ.get("AWS_REGION", "us-east-1"),
)
body = json.dumps({
"inputText": "the quick brown fox jumped over the lazy dog",
})
resp = client.invoke_model(
modelId=os.environ.get("BEDROCK_EMBED_MODEL", "amazon.titan-embed-text-v2:0"),
body=body,
accept="application/json",
contentType="application/json",
)
vector = json.loads(resp["body"].read())["embedding"] # list[float], len 1024
Batch embedding¶
Every SDK above takes a list; batch requests amortize latency. Azure OpenAI accepts up to 2,048 inputs per call with a 300k-token aggregate cap (source).
# Azure — batch shape
resp = client.embeddings.create(
model=os.environ["AZURE_EMBED_DEPLOYMENT"],
input=["chunk 1 text", "chunk 2 text", "chunk 3 text"],
)
vectors = [d.embedding for d in resp.data]
Same shape on Vertex (contents=[...]) and Bedrock (loop invoke_model — Titan doesn't expose a batch input).
A common signature¶
The app doesn't call these directly — it goes through the tiny helper in examples/rag/service/embed.py, which exposes:
CHIRON_PROVIDER picks which per-cloud path runs. Everything downstream — Store, Retrieve, Generate — is provider-agnostic once you have the vectors.