Skip to content

Store — native

Each cloud has a purpose-built vector store. This page shows the create-index / upload / query surface per provider. If you want a store you can also run on your laptop, jump to Store — pgvector.

Official docs verified 2026-08-08

The shape

The three surfaces differ meaningfully:

Provider Store class Index-time API Query-time API
Azure SearchClient (azure-search-documents) upload_documents(docs) where each doc has a Collection(Edm.Single) vector field search(vector_queries=[VectorizedQuery(...)])
GCP MatchingEngineIndexEndpoint (google-cloud-aiplatform) upsert_datapoints(datapoints=[IndexDatapoint(...)]) on the underlying MatchingEngineIndex find_neighbors(deployed_index_id, queries=[[...]])
AWS bedrock-agent-runtime (retrieve) — an S3 → KB pipeline ingests + embeds documents outside the app (out-of-band S3 sync) retrieve(knowledgeBaseId=…, retrievalQuery={"text": …})

Create the index once, then upload + query as a normal application. VectorizedQuery sends an already-embedded query vector; there's also VectorizableTextQuery if you want Azure AI Search to embed for you.

Create index

from azure.identity import DefaultAzureCredential
from azure.search.documents.indexes import SearchIndexClient
from azure.search.documents.indexes.models import (
    SearchIndex, SearchField, SimpleField, SearchableField,
    SearchFieldDataType, VectorSearch, HnswAlgorithmConfiguration,
    VectorSearchProfile,
)

fields = [
    SimpleField(name="id", type=SearchFieldDataType.String, key=True),
    SearchableField(name="text", type=SearchFieldDataType.String),
    SearchField(
        name="text_vector",
        type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
        searchable=True,
        vector_search_dimensions=1536,           # match text-embedding-3-small
        vector_search_profile_name="hnsw",
    ),
]
vector_search = VectorSearch(
    algorithms=[HnswAlgorithmConfiguration(name="hnsw-algo")],
    profiles=[VectorSearchProfile(name="hnsw", algorithm_configuration_name="hnsw-algo")],
)
index = SearchIndex(name="chiron-rag", fields=fields, vector_search=vector_search)

client = SearchIndexClient(endpoint=endpoint, credential=DefaultAzureCredential())
client.create_or_update_index(index)

Upload + query

from azure.search.documents import SearchClient
from azure.search.documents.models import VectorizedQuery

sc = SearchClient(endpoint=endpoint, index_name="chiron-rag",
                  credential=DefaultAzureCredential())

sc.upload_documents(documents=[
    {"id": "doc-1", "text": "chunk text here", "text_vector": vector},   # from embed.py
])

vq = VectorizedQuery(vector=query_vec, k_nearest_neighbors=5, fields="text_vector")
hits = sc.search(vector_queries=[vq], select=["id", "text"], top=5)
for h in hits:
    print(h["id"], h["text"])

Vertex separates the index (where vectors live) from the index endpoint (queryable HTTP surface). You upsert to the index; you query the endpoint.

from google.cloud import aiplatform

aiplatform.init(project=project, location=region)

index = aiplatform.MatchingEngineIndex(index_name=index_resource)
index.upsert_datapoints(datapoints=[
    aiplatform.matching_engine.matching_engine_index.IndexDatapoint(
        datapoint_id="doc-1",
        feature_vector=vector,                 # from embed.py
    ),
])

endpoint = aiplatform.MatchingEngineIndexEndpoint(
    index_endpoint_name=endpoint_resource,
)
matches = endpoint.find_neighbors(
    deployed_index_id="chiron_rag_deployed",
    queries=[query_vec],
    num_neighbors=5,
)
for m in matches[0]:
    print(m.id, m.distance)

AWS — Bedrock Knowledge Bases

Bedrock KB is the batteries-included path — you pipe documents into S3, a data-source connector chunks + embeds them into an OpenSearch Serverless index (or Aurora, or S3 vectors), and the app queries via bedrock-agent-runtime. No index-time SDK inside the app itself.

Query

import os, boto3

client = boto3.client(
    "bedrock-agent-runtime",
    region_name=os.environ.get("AWS_REGION", "us-east-1"),
)

resp = client.retrieve(
    knowledgeBaseId=os.environ["BEDROCK_KB_ID"],
    retrievalQuery={"text": query},
    retrievalConfiguration={
        "vectorSearchConfiguration": {"numberOfResults": 5},
    },
)
for r in resp["retrievalResults"]:
    print(r["content"]["text"], r.get("score"))

The one-shot retrieve_and_generate variant fuses retrieval + a Bedrock model call into one API — useful for the fastest path, but it hides the assembled prompt. This RAG recipe keeps retrieve and generate visible so generate.md can reuse the E2 chat service unchanged.

Trade-offs

Concern Azure AI Search Vertex Vector Search Bedrock KB
Semantic ranker built in ✅ via hybrid ✅ (LLM re-rank)
BYO chunker ✅ ✅ ⚠️ (default chunker via S3 pipeline)
Portability off cloud ❌ ❌ ❌
Same infra also on laptop ❌ ❌ ❌

The last row is what motivates the pgvector path — same schema you run on Cloud SQL / Azure Postgres / RDS also runs against docker run pgvector/pgvector:pg16.