RAG end-to-end¶
The second recipe on top of Foundations and Call an LLM. Same tri-parity discipline, one more step: retrieval-augmented generation.
The point of E3 is that the pieces compose. The E2 chat service is the generation step here — E3 does not rewrite it. Everything the reader already learned about Azure OpenAI / Vertex Gemini / Bedrock stays intact; RAG just wraps retrieve around generate.
The pipeline¶
┌──────────┐ ┌──────────┐ ┌────────────┐ ┌──────────────┐
│ embed │──►│ store │──►│ retrieve │──►│ generate │
│ (model) │ │ (vector) │ │ (top-K) │ │ (E2 chat) │
└──────────┘ └──────────┘ └────────────┘ └──────────────┘
per-cloud dual-store dual-store REUSED verbatim
flagship (native + pgvector) from Call an LLM
Two storage paths per cloud¶
Every provider gets both storage backends. Pick whichever fits.
| Cloud | Native vector store | Provider-agnostic |
|---|---|---|
| Azure | Azure AI Search (azure-search-documents) |
pgvector on Azure Database for PostgreSQL Flexible Server |
| GCP | Vertex AI Vector Search | pgvector on Cloud SQL for PostgreSQL |
| AWS | Bedrock Knowledge Bases (bedrock-agent-runtime) |
pgvector on RDS for PostgreSQL |
The native side is where you sign up for the cloud's fully-managed features (semantic ranking, hybrid, reranking). The pgvector side is where you own the store — same Python, same SQL, same schema on all three clouds. Pick the row where the trade-off lands right for your team.
Docs verified 2026-08-08¶
Every model id + SDK below was verified against these live pages:
- Azure OpenAI embeddings (
text-embedding-3-small,text-embedding-3-large) — learn.microsoft.com/…/openai/how-to/embeddings - Azure AI Search vector quickstart (
azure-search-documents,VectorizedQuery) — learn.microsoft.com/…/search/search-get-started-vector - Vertex AI embeddings (
gemini-embedding-001,text-embedding-005) — ai.google.dev/gemini-api/docs/embeddings - Bedrock Titan embeddings (
amazon.titan-embed-text-v2:0, 1024 dims) — docs.aws.amazon.com/bedrock/…/titan-embedding-models - Bedrock Knowledge Bases retrieve API (
bedrock-agent-runtime,retrieve_and_generate) — docs.aws.amazon.com/bedrock/…/kb-test-config - pgvector — github.com/pgvector/pgvector
Pages¶
- Embed — the embedding model per provider + a common Python signature the app uses.
- Store — native — Azure AI Search / Vertex Vector Search / Bedrock KB.
- Store — pgvector — one SQL schema, three managed Postgres flavors.
- Retrieve — embed the query, top-K similarity.
- Generate — feed retrieved chunks into the E2 chat service. This is a reuse chapter, not a rewrite.
- Deploy — Terraform composing the E1 baseline + provider vector store + Postgres, Helm reusing the E2 chart.
Validate-only¶
Every gate on this batch is a schema/lint check — terraform validate, helm lint, mkdocs build --strict, python -m compileall, plus Chiron's completeness-check + leak-scan. Live LLM + live vector-store calls are laptop operations; the "verify" block on each page shows the one-liner. Phase-3 billed apply lands separately.