Skip to content

Retrieve

Embed the user query with the same model that embedded the corpus, then top-K similarity search against whichever store you chose. Two lines when the pieces from the earlier pages are in scope.

Docs cross-reference

Uses Embed + Store — native / Store — pgvector. No new SDKs on this page.

The pattern

from embed import embed_texts           # from docs/rag/embed.md
from store_pgvector import search_similar   # or store_azure, store_gcp, store_aws

def retrieve(query: str, k: int = 5) -> list[dict]:
    [query_vec] = embed_texts([query])
    return search_similar(query_vec, k=k)

That's the entire retrieve step. Whatever search_similar returns — a list of {id, text, similarity, …} dicts — is what Generate feeds the model.

Native paths — the same function, different backend

from azure.search.documents import SearchClient
from azure.search.documents.models import VectorizedQuery

def search_similar_azure(sc: SearchClient, query_vec, k=5):
    vq = VectorizedQuery(
        vector=query_vec,
        k_nearest_neighbors=k,
        fields="text_vector",
    )
    hits = sc.search(vector_queries=[vq], select=["id", "text"], top=k)
    return [{"id": h["id"], "text": h["text"], "score": h["@search.score"]} for h in hits]
from google.cloud import aiplatform

def search_similar_gcp(endpoint: aiplatform.MatchingEngineIndexEndpoint, query_vec, k=5):
    matches = endpoint.find_neighbors(
        deployed_index_id="chiron_rag_deployed",
        queries=[query_vec],
        num_neighbors=k,
    )
    return [{"id": m.id, "score": 1 - m.distance} for m in matches[0]]

The Vertex response carries datapoint ids only — you fetch the full text from your source of record keyed by id. That trade-off is characteristic of the Matching Engine: fast, but not a document store.

Bedrock KB returns the chunk text alongside the score, so the shape matches pgvector's search_similar directly.

import os, boto3

_client = boto3.client("bedrock-agent-runtime", region_name=os.environ["AWS_REGION"])

def search_similar_aws(query: str, k: int = 5):
    # Bedrock KB embeds the query itself — no embed_texts call needed.
    resp = _client.retrieve(
        knowledgeBaseId=os.environ["BEDROCK_KB_ID"],
        retrievalQuery={"text": query},
        retrievalConfiguration={
            "vectorSearchConfiguration": {"numberOfResults": k},
        },
    )
    return [
        {"id": r["location"].get("s3Location", {}).get("uri", "?"),
         "text": r["content"]["text"],
         "score": r.get("score")}
        for r in resp["retrievalResults"]
    ]

pgvector — the portable path

Already documented on Store — pgvector. Same signature as the native adapters above.

Top-K choice

k=5 is a sensible default for chat context (the E2 chat models comfortably accept 5×1KB chunks). Turn it up for reranking pipelines (k=25 → rerank → top 5) or turn it down for factual short answers. The rest of the pipeline doesn't care what value you pick.