Retrieve¶
Embed the user query with the same model that embedded the corpus, then top-K similarity search against whichever store you chose. Two lines when the pieces from the earlier pages are in scope.
Docs cross-reference
Uses Embed + Store — native / Store — pgvector. No new SDKs on this page.
The pattern¶
from embed import embed_texts # from docs/rag/embed.md
from store_pgvector import search_similar # or store_azure, store_gcp, store_aws
def retrieve(query: str, k: int = 5) -> list[dict]:
[query_vec] = embed_texts([query])
return search_similar(query_vec, k=k)
That's the entire retrieve step. Whatever search_similar returns — a list of {id, text, similarity, …} dicts — is what Generate feeds the model.
Native paths — the same function, different backend¶
from azure.search.documents import SearchClient
from azure.search.documents.models import VectorizedQuery
def search_similar_azure(sc: SearchClient, query_vec, k=5):
vq = VectorizedQuery(
vector=query_vec,
k_nearest_neighbors=k,
fields="text_vector",
)
hits = sc.search(vector_queries=[vq], select=["id", "text"], top=k)
return [{"id": h["id"], "text": h["text"], "score": h["@search.score"]} for h in hits]
from google.cloud import aiplatform
def search_similar_gcp(endpoint: aiplatform.MatchingEngineIndexEndpoint, query_vec, k=5):
matches = endpoint.find_neighbors(
deployed_index_id="chiron_rag_deployed",
queries=[query_vec],
num_neighbors=k,
)
return [{"id": m.id, "score": 1 - m.distance} for m in matches[0]]
The Vertex response carries datapoint ids only — you fetch the full text from your source of record keyed by id. That trade-off is characteristic of the Matching Engine: fast, but not a document store.
Bedrock KB returns the chunk text alongside the score, so the shape matches pgvector's search_similar directly.
import os, boto3
_client = boto3.client("bedrock-agent-runtime", region_name=os.environ["AWS_REGION"])
def search_similar_aws(query: str, k: int = 5):
# Bedrock KB embeds the query itself — no embed_texts call needed.
resp = _client.retrieve(
knowledgeBaseId=os.environ["BEDROCK_KB_ID"],
retrievalQuery={"text": query},
retrievalConfiguration={
"vectorSearchConfiguration": {"numberOfResults": k},
},
)
return [
{"id": r["location"].get("s3Location", {}).get("uri", "?"),
"text": r["content"]["text"],
"score": r.get("score")}
for r in resp["retrievalResults"]
]
pgvector — the portable path¶
Already documented on Store — pgvector. Same signature as the native adapters above.
Top-K choice¶
k=5 is a sensible default for chat context (the E2 chat models comfortably accept 5×1KB chunks). Turn it up for reranking pipelines (k=25 → rerank → top 5) or turn it down for factual short answers. The rest of the pipeline doesn't care what value you pick.