Store — native¶
Each cloud has a purpose-built vector store. This page shows the create-index / upload / query surface per provider. If you want a store you can also run on your laptop, jump to Store — pgvector.
Official docs verified 2026-08-08
- Azure AI Search vector quickstart (
azure-search-documents): learn.microsoft.com/…/search/search-get-started-vector - Vertex AI Vector Search: cloud.google.com/vertex-ai/docs/vector-search/overview
- Bedrock Knowledge Bases (
bedrock-agent-runtime): docs.aws.amazon.com/bedrock/…/kb-test-config
The shape¶
The three surfaces differ meaningfully:
| Provider | Store class | Index-time API | Query-time API |
|---|---|---|---|
| Azure | SearchClient (azure-search-documents) |
upload_documents(docs) where each doc has a Collection(Edm.Single) vector field |
search(vector_queries=[VectorizedQuery(...)]) |
| GCP | MatchingEngineIndexEndpoint (google-cloud-aiplatform) |
upsert_datapoints(datapoints=[IndexDatapoint(...)]) on the underlying MatchingEngineIndex |
find_neighbors(deployed_index_id, queries=[[...]]) |
| AWS | bedrock-agent-runtime (retrieve) — an S3 → KB pipeline ingests + embeds documents outside the app |
(out-of-band S3 sync) | retrieve(knowledgeBaseId=…, retrievalQuery={"text": …}) |
Azure — Azure AI Search¶
Create the index once, then upload + query as a normal application. VectorizedQuery sends an already-embedded query vector; there's also VectorizableTextQuery if you want Azure AI Search to embed for you.
Create index¶
from azure.identity import DefaultAzureCredential
from azure.search.documents.indexes import SearchIndexClient
from azure.search.documents.indexes.models import (
SearchIndex, SearchField, SimpleField, SearchableField,
SearchFieldDataType, VectorSearch, HnswAlgorithmConfiguration,
VectorSearchProfile,
)
fields = [
SimpleField(name="id", type=SearchFieldDataType.String, key=True),
SearchableField(name="text", type=SearchFieldDataType.String),
SearchField(
name="text_vector",
type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
searchable=True,
vector_search_dimensions=1536, # match text-embedding-3-small
vector_search_profile_name="hnsw",
),
]
vector_search = VectorSearch(
algorithms=[HnswAlgorithmConfiguration(name="hnsw-algo")],
profiles=[VectorSearchProfile(name="hnsw", algorithm_configuration_name="hnsw-algo")],
)
index = SearchIndex(name="chiron-rag", fields=fields, vector_search=vector_search)
client = SearchIndexClient(endpoint=endpoint, credential=DefaultAzureCredential())
client.create_or_update_index(index)
Upload + query¶
from azure.search.documents import SearchClient
from azure.search.documents.models import VectorizedQuery
sc = SearchClient(endpoint=endpoint, index_name="chiron-rag",
credential=DefaultAzureCredential())
sc.upload_documents(documents=[
{"id": "doc-1", "text": "chunk text here", "text_vector": vector}, # from embed.py
])
vq = VectorizedQuery(vector=query_vec, k_nearest_neighbors=5, fields="text_vector")
hits = sc.search(vector_queries=[vq], select=["id", "text"], top=5)
for h in hits:
print(h["id"], h["text"])
GCP — Vertex AI Vector Search¶
Vertex separates the index (where vectors live) from the index endpoint (queryable HTTP surface). You upsert to the index; you query the endpoint.
from google.cloud import aiplatform
aiplatform.init(project=project, location=region)
index = aiplatform.MatchingEngineIndex(index_name=index_resource)
index.upsert_datapoints(datapoints=[
aiplatform.matching_engine.matching_engine_index.IndexDatapoint(
datapoint_id="doc-1",
feature_vector=vector, # from embed.py
),
])
endpoint = aiplatform.MatchingEngineIndexEndpoint(
index_endpoint_name=endpoint_resource,
)
matches = endpoint.find_neighbors(
deployed_index_id="chiron_rag_deployed",
queries=[query_vec],
num_neighbors=5,
)
for m in matches[0]:
print(m.id, m.distance)
AWS — Bedrock Knowledge Bases¶
Bedrock KB is the batteries-included path — you pipe documents into S3, a data-source connector chunks + embeds them into an OpenSearch Serverless index (or Aurora, or S3 vectors), and the app queries via bedrock-agent-runtime. No index-time SDK inside the app itself.
Query¶
import os, boto3
client = boto3.client(
"bedrock-agent-runtime",
region_name=os.environ.get("AWS_REGION", "us-east-1"),
)
resp = client.retrieve(
knowledgeBaseId=os.environ["BEDROCK_KB_ID"],
retrievalQuery={"text": query},
retrievalConfiguration={
"vectorSearchConfiguration": {"numberOfResults": 5},
},
)
for r in resp["retrievalResults"]:
print(r["content"]["text"], r.get("score"))
The one-shot retrieve_and_generate variant fuses retrieval + a Bedrock model call into one API — useful for the fastest path, but it hides the assembled prompt. This RAG recipe keeps retrieve and generate visible so generate.md can reuse the E2 chat service unchanged.
Trade-offs¶
| Concern | Azure AI Search | Vertex Vector Search | Bedrock KB |
|---|---|---|---|
| Semantic ranker built in | ✅ | via hybrid | ✅ (LLM re-rank) |
| BYO chunker | ✅ | ✅ | ⚠️ (default chunker via S3 pipeline) |
| Portability off cloud | ❌ | ❌ | ❌ |
| Same infra also on laptop | ❌ | ❌ | ❌ |
The last row is what motivates the pgvector path — same schema you run on Cloud SQL / Azure Postgres / RDS also runs against docker run pgvector/pgvector:pg16.