Call an LLM¶
The first working feature on top of the E1 Foundations. One chat service, containerized, deployed both ways per provider, with two code shapes per provider:
- The native flagship on each cloud — Azure OpenAI (GPT), Vertex AI (Gemini), Bedrock (Amazon Nova).
- The Claude through-line — the same Anthropic-authored Python SDK usable through per-cloud clients (Bedrock, Vertex, Foundry). One SDK vocabulary; three cloud-native auth surfaces.
Every code snippet on these pages is verified against the current provider docs (URLs + verified-on dates land at the top of each page). No model IDs are recalled from memory.
The through-line — why Claude appears on all three clouds¶
Anthropic ships a per-cloud Python client so you can point the same API surface at any of the three:
| Cloud | pip extras | Client class | Import |
|---|---|---|---|
| AWS Bedrock | anthropic[bedrock] |
AnthropicBedrock |
from anthropic import AnthropicBedrock |
| GCP Vertex | anthropic[vertex] |
AnthropicVertex |
from anthropic import AnthropicVertex |
| Azure Foundry | anthropic + azure-identity |
AnthropicFoundry |
from anthropic import AnthropicFoundry |
The application code is the same everywhere. The AUTH surface and MODEL ID differ per cloud. That's the load-bearing point of Chiron: the choice of cloud is a scoped decision, not a rewrite.
The service — one container, deployed six ways¶
The reader-facing artifact is the same each time: an HTTP service that takes a prompt and streams a chat response. It compiles per-provider:
- Kubernetes (AKS / GKE / EKS) — the E1 baseline provisions the cluster + a container registry; the chart extends the E1 chart with the chat service and the env-vars the chosen SDK needs.
- Serverless (Container Apps / Cloud Run / App Runner) — the same image, deployed as a scale-to-zero HTTP endpoint.
Terraform for each provider composes the E1 baseline as a module and layers the model access (IAM role, endpoint URL env vars) on top. You'll never see the E1 plumbing duplicated.
Pages¶
- Chat — the request/response shape per provider, both native + Claude.
- Streaming — turn on incremental output.
- Tools (function calling) — the native flagship's tool-use surface per provider.
- Deploy — Terraform + Helm reusing the E1 baseline; deploys both K8s + serverless.
Validate-only, no live LLM calls in CI¶
Everything on these pages terraform validates, tflints, helm lints, and Python-imports without money changing hands. Live LLM calls are a laptop-only operation — the "verify" block on each page shows the one-liner to run against your own credentials.