Deploy the RAG service¶
Same shape as Call an LLM → Deploy. The Terraform module composes the E1 baseline and the E2 call-an-llm module; the Helm chart adds a database URL to the E2 chart's env. No baseline or E2 code is duplicated.
Official docs verified 2026-08-08
- Azure Database for PostgreSQL Flexible Server (Terraform): registry.terraform.io/providers/hashicorp/azurerm/latest/docs/resources/postgresql_flexible_server
- Cloud SQL for PostgreSQL (Terraform): registry.terraform.io/providers/hashicorp/google/latest/docs/resources/sql_database_instance
- RDS for PostgreSQL (Terraform): registry.terraform.io/providers/hashicorp/aws/latest/docs/resources/db_instance
Layout¶
examples/rag/
service/ # extends the E2 chat service with retrieval
embed.py # provider-picking embed helper
store_pgvector.py # pgvector Python client (managed-Postgres agnostic)
store_native.py # per-provider native adapter
retrieve.py # embed → similarity
rag.py # retrieve + E2 chat compose
main.py # FastAPI /rag + /rag-claude
Dockerfile # same base as E2; adds pgvector + azure-search-documents
requirements.txt
chart/ # extends the E2 chart with DATABASE_URL + KB_ID env
Chart.yaml
values.yaml
templates/
deployment.yaml
service.yaml
terraform/
azure/ (module "baseline" + "call_an_llm" + azurerm_postgresql_flexible_server)
gcp/ (module "baseline" + "call_an_llm" + google_sql_database_instance)
aws/ (module "baseline" + "call_an_llm" + aws_db_instance)
Terraform — compose, don't duplicate¶
Each provider's module pulls in both upstream modules:
module "baseline" {
source = "../../../foundations/azure"
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/azure"
name = var.name
region = var.region
compute = var.compute
azure_openai_base_url = var.azure_openai_base_url
azure_openai_deployment = var.azure_openai_deployment
foundry_claude_deployment = var.foundry_claude_deployment
}
# Add the RAG-only piece — a managed Postgres with pgvector.
resource "azurerm_postgresql_flexible_server" "chiron_rag" {
name = "${var.name}-pg"
resource_group_name = module.baseline.resource_group_name
location = var.region
administrator_login = var.pg_admin_login
administrator_password = var.pg_admin_password
sku_name = "B_Standard_B1ms"
version = "16"
storage_mb = 32768
}
resource "azurerm_postgresql_flexible_server_configuration" "vector_ext" {
name = "azure.extensions"
server_id = azurerm_postgresql_flexible_server.chiron_rag.id
value = "VECTOR"
}
module "baseline" {
source = "../../../foundations/gcp"
project_id = var.project_id
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/gcp"
project_id = var.project_id
name = var.name
region = var.region
compute = var.compute
}
resource "google_sql_database_instance" "chiron_rag" {
name = "${var.name}-pg"
database_version = "POSTGRES_16"
region = var.region
settings {
tier = "db-custom-1-3840"
database_flags {
name = "cloudsql.iam_authentication"
value = "on"
}
}
deletion_protection = false
}
module "baseline" {
source = "../../../foundations/aws"
name = var.name
region = var.region
compute = var.compute
}
module "call_an_llm" {
source = "../../../call-an-llm/terraform/aws"
name = var.name
region = var.region
compute = var.compute
}
resource "aws_db_instance" "chiron_rag" {
identifier = "${var.name}-pg"
engine = "postgres"
engine_version = "16.4"
instance_class = "db.t4g.micro"
allocated_storage = 20
db_name = "chiron_rag"
username = var.pg_admin_login
password = var.pg_admin_password
publicly_accessible = false
skip_final_snapshot = true
apply_immediately = true
backup_retention_period = 0
}
Helm — one small extension of the E2 chart¶
The E3 chart at examples/rag/chart/ templates the same deployment shape as E2 plus one required env var (DATABASE_URL) and one optional env var (BEDROCK_KB_ID). The chart is installed alongside — or instead of — the E2 chart depending on whether you want the plain /chat endpoint too.
helm upgrade --install rag ./examples/rag/chart \
--set image.repository="$(terraform -chdir=examples/rag/terraform/azure output -raw container_registry_login_server)/rag" \
--set image.tag="v1" \
--set env.CHIRON_PROVIDER="azure" \
--set env.AZURE_OPENAI_BASE_URL="https://<res>.openai.azure.com/openai/v1/" \
--set env.AZURE_OPENAI_DEPLOYMENT="gpt-5" \
--set env.AZURE_EMBED_DEPLOYMENT="text-embedding-3-small" \
--set env.DATABASE_URL="postgresql://…"
Verify (validate-only)¶
# Terraform
for p in azure gcp aws; do
( cd examples/rag/terraform/$p && terraform init -backend=false && terraform validate )
done
# Helm
helm lint examples/rag/chart
helm template rag examples/rag/chart > /dev/null
# Python — compile smoke
python -m compileall examples/rag/service
Live apply¶
Same command inverted for teardown:
E3 is validate-only — live RAG requires real cloud resources + a real corpus. Phase-3 live-apply lands separately.