Skip to content

Embed

Turn text into a vector. Every RAG pipeline starts here; every provider has a per-cloud SDK path that produces a comparable-shape response.

Official docs verified 2026-08-08

Model IDs

Provider Model ID (default) Dimensions Notes
Azure OpenAI text-embedding-3-small 1536 -large variant: 3072. Uses your DEPLOYMENT name.
Vertex AI gemini-embedding-001 3072 (default) Older stable IDs: text-embedding-005, text-multilingual-embedding-002.
Bedrock amazon.titan-embed-text-v2:0 1024 (default), 512, 256 Older stable ID: amazon.titan-embed-text-v1 (1536 dims).

The dimension choice matters — the store schema has to match. All examples below hold to the default row of the table.

Native flagship

Uses the same openai package + openai/v1/ base URL as Call an LLM → Chat. model= is the DEPLOYMENT name of your embedding model.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AZURE_OPENAI_API_KEY"],
    base_url=os.environ["AZURE_OPENAI_BASE_URL"],   # …/openai/v1/
)

resp = client.embeddings.create(
    model=os.environ.get("AZURE_EMBED_DEPLOYMENT", "text-embedding-3-small"),
    input="the quick brown fox jumped over the lazy dog",
)
vector = resp.data[0].embedding                     # list[float], len 1536
import os
from google import genai

client = genai.Client(
    vertexai=True,
    project=os.environ["GOOGLE_CLOUD_PROJECT"],
    location=os.environ.get("GOOGLE_CLOUD_LOCATION", "us-central1"),
)

resp = client.models.embed_content(
    model=os.environ.get("GEMINI_EMBED_MODEL", "gemini-embedding-001"),
    contents="the quick brown fox jumped over the lazy dog",
)
vector = resp.embeddings[0].values                  # list[float]

Bedrock embeddings are invoke_model (not Converse — Converse doesn't cover embeddings, per the Titan embed docs).

import json, os, boto3

client = boto3.client(
    "bedrock-runtime",
    region_name=os.environ.get("AWS_REGION", "us-east-1"),
)
body = json.dumps({
    "inputText": "the quick brown fox jumped over the lazy dog",
})
resp = client.invoke_model(
    modelId=os.environ.get("BEDROCK_EMBED_MODEL", "amazon.titan-embed-text-v2:0"),
    body=body,
    accept="application/json",
    contentType="application/json",
)
vector = json.loads(resp["body"].read())["embedding"]   # list[float], len 1024

Batch embedding

Every SDK above takes a list; batch requests amortize latency. Azure OpenAI accepts up to 2,048 inputs per call with a 300k-token aggregate cap (source).

# Azure — batch shape
resp = client.embeddings.create(
    model=os.environ["AZURE_EMBED_DEPLOYMENT"],
    input=["chunk 1 text", "chunk 2 text", "chunk 3 text"],
)
vectors = [d.embedding for d in resp.data]

Same shape on Vertex (contents=[...]) and Bedrock (loop invoke_model — Titan doesn't expose a batch input).

A common signature

The app doesn't call these directly — it goes through the tiny helper in examples/rag/service/embed.py, which exposes:

def embed_texts(texts: list[str]) -> list[list[float]]: ...

CHIRON_PROVIDER picks which per-cloud path runs. Everything downstream — Store, Retrieve, Generate — is provider-agnostic once you have the vectors.