Skip to content

RAG end-to-end

The second recipe on top of Foundations and Call an LLM. Same tri-parity discipline, one more step: retrieval-augmented generation.

The point of E3 is that the pieces compose. The E2 chat service is the generation step here — E3 does not rewrite it. Everything the reader already learned about Azure OpenAI / Vertex Gemini / Bedrock stays intact; RAG just wraps retrieve around generate.

The pipeline

     ┌──────────┐   ┌──────────┐   ┌────────────┐   ┌──────────────┐
     │  embed   │──►│  store   │──►│  retrieve  │──►│  generate    │
     │ (model)  │   │ (vector) │   │  (top-K)   │   │  (E2 chat)   │
     └──────────┘   └──────────┘   └────────────┘   └──────────────┘
       per-cloud   dual-store       dual-store         REUSED verbatim
       flagship    (native + pgvector)                 from Call an LLM

Two storage paths per cloud

Every provider gets both storage backends. Pick whichever fits.

Cloud Native vector store Provider-agnostic
Azure Azure AI Search (azure-search-documents) pgvector on Azure Database for PostgreSQL Flexible Server
GCP Vertex AI Vector Search pgvector on Cloud SQL for PostgreSQL
AWS Bedrock Knowledge Bases (bedrock-agent-runtime) pgvector on RDS for PostgreSQL

The native side is where you sign up for the cloud's fully-managed features (semantic ranking, hybrid, reranking). The pgvector side is where you own the store — same Python, same SQL, same schema on all three clouds. Pick the row where the trade-off lands right for your team.

Docs verified 2026-08-08

Every model id + SDK below was verified against these live pages:

Pages

  • Embed — the embedding model per provider + a common Python signature the app uses.
  • Store — native — Azure AI Search / Vertex Vector Search / Bedrock KB.
  • Store — pgvector — one SQL schema, three managed Postgres flavors.
  • Retrieve — embed the query, top-K similarity.
  • Generate — feed retrieved chunks into the E2 chat service. This is a reuse chapter, not a rewrite.
  • Deploy — Terraform composing the E1 baseline + provider vector store + Postgres, Helm reusing the E2 chart.

Validate-only

Every gate on this batch is a schema/lint check — terraform validate, helm lint, mkdocs build --strict, python -m compileall, plus Chiron's completeness-check + leak-scan. Live LLM + live vector-store calls are laptop operations; the "verify" block on each page shows the one-liner. Phase-3 billed apply lands separately.