Generate¶
This page reuses Call an LLM → Chat verbatim. Nothing new to learn here — RAG is a prompt-composition trick on top of the E2 chat surface. The load-bearing observation is that both the native flagship and the Claude through-line already know how to accept the assembled prompt.
The composition¶
from chat_native import azure_chat, gcp_chat, aws_chat # from E2 examples/call-an-llm/service/
from retrieve import retrieve
_NATIVE = {"azure": azure_chat, "gcp": gcp_chat, "aws": aws_chat}
RAG_TEMPLATE = """Use the following context to answer the user's question.
If the context is insufficient, say so; do not invent facts.
Context
-------
{context}
Question
--------
{question}
"""
def rag_answer(provider: str, question: str, k: int = 5) -> dict:
chunks = retrieve(question, k=k)
prompt = RAG_TEMPLATE.format(
context="\n\n".join(f"[{c['id']}] {c['text']}" for c in chunks),
question=question,
)
reply = _NATIVE[provider](prompt)
return {"reply": reply, "sources": [c["id"] for c in chunks]}
That's the entire generate step. The E2 chat functions are imported unmodified — no per-cloud RAG branch, no shim, no wrapper.
Claude through-line composes identically¶
Swap _NATIVE for _CLAUDE:
from chat_claude import foundry_claude, vertex_claude, bedrock_claude
_CLAUDE = {"azure": foundry_claude, "gcp": vertex_claude, "aws": bedrock_claude}
The application code around it doesn't change — the RAG prompt goes in as one string, the reply comes out. The through-line's whole point is that the app-level shape is stable across clouds.
Citations¶
The sources list in the return value is the citation trail. On Azure AI Search and pgvector, id is your own document key. On Vertex Vector Search, id is the datapoint id (fetch the source text from your record store). On Bedrock KB, id is the S3 URI of the origin document.
For a citation-preserving prompt shape (embed the citation tag next to each chunk so the model can reference it), the template above already includes [{id}] inline — the E2 models pick it up reliably. If you need structured citations at the model layer, Tools is the extension point.
Streaming answers¶
Wire retrieve + the streaming versions from Call an LLM → Streaming — same composition, same one-line swap. Chunks come back token-by-token; the sources header still lands before the first token if you serialize it up front in the response.
The one thing to NOT do¶
Don't rewrite the chat call inside a RAG module. The whole point of composition is that the two recipes stay orthogonal — future E-series recipes (evals, agents) also compose the E2 chat surface, and if you fork it here, you have to keep three near-identical copies in sync forever.