Skip to content

Call an LLM

The first working feature on top of the E1 Foundations. One chat service, containerized, deployed both ways per provider, with two code shapes per provider:

  1. The native flagship on each cloud — Azure OpenAI (GPT), Vertex AI (Gemini), Bedrock (Amazon Nova).
  2. The Claude through-line — the same Anthropic-authored Python SDK usable through per-cloud clients (Bedrock, Vertex, Foundry). One SDK vocabulary; three cloud-native auth surfaces.

Every code snippet on these pages is verified against the current provider docs (URLs + verified-on dates land at the top of each page). No model IDs are recalled from memory.

The through-line — why Claude appears on all three clouds

Anthropic ships a per-cloud Python client so you can point the same API surface at any of the three:

Cloud pip extras Client class Import
AWS Bedrock anthropic[bedrock] AnthropicBedrock from anthropic import AnthropicBedrock
GCP Vertex anthropic[vertex] AnthropicVertex from anthropic import AnthropicVertex
Azure Foundry anthropic + azure-identity AnthropicFoundry from anthropic import AnthropicFoundry

The application code is the same everywhere. The AUTH surface and MODEL ID differ per cloud. That's the load-bearing point of Chiron: the choice of cloud is a scoped decision, not a rewrite.

The service — one container, deployed six ways

The reader-facing artifact is the same each time: an HTTP service that takes a prompt and streams a chat response. It compiles per-provider:

  • Kubernetes (AKS / GKE / EKS) — the E1 baseline provisions the cluster + a container registry; the chart extends the E1 chart with the chat service and the env-vars the chosen SDK needs.
  • Serverless (Container Apps / Cloud Run / App Runner) — the same image, deployed as a scale-to-zero HTTP endpoint.

Terraform for each provider composes the E1 baseline as a module and layers the model access (IAM role, endpoint URL env vars) on top. You'll never see the E1 plumbing duplicated.

Pages

  • Chat — the request/response shape per provider, both native + Claude.
  • Streaming — turn on incremental output.
  • Tools (function calling) — the native flagship's tool-use surface per provider.
  • Deploy — Terraform + Helm reusing the E1 baseline; deploys both K8s + serverless.

Validate-only, no live LLM calls in CI

Everything on these pages terraform validates, tflints, helm lints, and Python-imports without money changing hands. Live LLM calls are a laptop-only operation — the "verify" block on each page shows the one-liner to run against your own credentials.