Skip to content

Deploy the RAG service

Same shape as Call an LLM → Deploy. The Terraform module composes the E1 baseline and the E2 call-an-llm module; the Helm chart adds a database URL to the E2 chart's env. No baseline or E2 code is duplicated.

Layout

examples/rag/
  service/                        # extends the E2 chat service with retrieval
    embed.py                      # provider-picking embed helper
    store_pgvector.py             # pgvector Python client (managed-Postgres agnostic)
    store_native.py               # per-provider native adapter
    retrieve.py                   # embed → similarity
    rag.py                        # retrieve + E2 chat compose
    main.py                       # FastAPI /rag + /rag-claude
    Dockerfile                    # same base as E2; adds pgvector + azure-search-documents
    requirements.txt
  chart/                          # extends the E2 chart with DATABASE_URL + KB_ID env
    Chart.yaml
    values.yaml
    templates/
      deployment.yaml
      service.yaml
  terraform/
    azure/     (module "baseline" + "call_an_llm" + azurerm_postgresql_flexible_server)
    gcp/       (module "baseline" + "call_an_llm" + google_sql_database_instance)
    aws/       (module "baseline" + "call_an_llm" + aws_db_instance)

Terraform — compose, don't duplicate

Each provider's module pulls in both upstream modules:

module "baseline" {
  source  = "../../../foundations/azure"
  name    = var.name
  region  = var.region
  compute = var.compute
}

module "call_an_llm" {
  source                     = "../../../call-an-llm/terraform/azure"
  name                       = var.name
  region                     = var.region
  compute                    = var.compute
  azure_openai_base_url      = var.azure_openai_base_url
  azure_openai_deployment    = var.azure_openai_deployment
  foundry_claude_deployment  = var.foundry_claude_deployment
}

# Add the RAG-only piece — a managed Postgres with pgvector.
resource "azurerm_postgresql_flexible_server" "chiron_rag" {
  name                = "${var.name}-pg"
  resource_group_name = module.baseline.resource_group_name
  location            = var.region
  administrator_login    = var.pg_admin_login
  administrator_password = var.pg_admin_password
  sku_name = "B_Standard_B1ms"
  version  = "16"
  storage_mb = 32768
}

resource "azurerm_postgresql_flexible_server_configuration" "vector_ext" {
  name      = "azure.extensions"
  server_id = azurerm_postgresql_flexible_server.chiron_rag.id
  value     = "VECTOR"
}
module "baseline" {
  source     = "../../../foundations/gcp"
  project_id = var.project_id
  name       = var.name
  region     = var.region
  compute    = var.compute
}

module "call_an_llm" {
  source     = "../../../call-an-llm/terraform/gcp"
  project_id = var.project_id
  name       = var.name
  region     = var.region
  compute    = var.compute
}

resource "google_sql_database_instance" "chiron_rag" {
  name             = "${var.name}-pg"
  database_version = "POSTGRES_16"
  region           = var.region

  settings {
    tier = "db-custom-1-3840"
    database_flags {
      name  = "cloudsql.iam_authentication"
      value = "on"
    }
  }
  deletion_protection = false
}
module "baseline" {
  source  = "../../../foundations/aws"
  name    = var.name
  region  = var.region
  compute = var.compute
}

module "call_an_llm" {
  source  = "../../../call-an-llm/terraform/aws"
  name    = var.name
  region  = var.region
  compute = var.compute
}

resource "aws_db_instance" "chiron_rag" {
  identifier              = "${var.name}-pg"
  engine                  = "postgres"
  engine_version          = "16.4"
  instance_class          = "db.t4g.micro"
  allocated_storage       = 20
  db_name                 = "chiron_rag"
  username                = var.pg_admin_login
  password                = var.pg_admin_password
  publicly_accessible     = false
  skip_final_snapshot     = true
  apply_immediately       = true
  backup_retention_period = 0
}

Helm — one small extension of the E2 chart

The E3 chart at examples/rag/chart/ templates the same deployment shape as E2 plus one required env var (DATABASE_URL) and one optional env var (BEDROCK_KB_ID). The chart is installed alongside — or instead of — the E2 chart depending on whether you want the plain /chat endpoint too.

helm upgrade --install rag ./examples/rag/chart \
  --set image.repository="$(terraform -chdir=examples/rag/terraform/azure output -raw container_registry_login_server)/rag" \
  --set image.tag="v1" \
  --set env.CHIRON_PROVIDER="azure" \
  --set env.AZURE_OPENAI_BASE_URL="https://<res>.openai.azure.com/openai/v1/" \
  --set env.AZURE_OPENAI_DEPLOYMENT="gpt-5" \
  --set env.AZURE_EMBED_DEPLOYMENT="text-embedding-3-small" \
  --set env.DATABASE_URL="postgresql://…"

Verify (validate-only)

# Terraform
for p in azure gcp aws; do
  ( cd examples/rag/terraform/$p && terraform init -backend=false && terraform validate )
done

# Helm
helm lint examples/rag/chart
helm template rag examples/rag/chart > /dev/null

# Python — compile smoke
python -m compileall examples/rag/service

Live apply

Same command inverted for teardown:

terraform destroy   # k8s and serverless

E3 is validate-only — live RAG requires real cloud resources + a real corpus. Phase-3 live-apply lands separately.