Billable units per service
Prices drift; units don't. This table maps every service Chiron uses to the exact metric each cloud invoices.
Live pricing (verified 2026-08-08)
Chat / embeddings / vision
| Service |
Billable units |
| Azure OpenAI (GPT / o-series / embeddings / vision) |
per 1M input tokens, per 1M output tokens, per 1M cached input tokens (reduced), per-image tokens on multimodal (counted into the input-token bucket) |
| Vertex Gemini (text / image-in / audio-in / video-in) |
per 1M input tokens, per 1M output tokens, per 1M cached input tokens (reduced) — a different per-1M rate applies above the 200k-token boundary on some models |
| Bedrock (Anthropic / Nova / others) |
per 1M input tokens and per 1M output tokens (invoice line item: "on-demand invocations") |
Vision input on all three is metered as tokens — the image is tokenized before the model reads it (Anthropic's rule is ⌈w/28⌉·⌈h/28⌉ visual tokens per image, capped by the model's tier — see Multimodal).
Image generation
| Service |
Billable units |
Azure OpenAI (gpt-image-*, historically DALL-E) |
per image, priced by size × quality × output_format |
Vertex Imagen (imagen-4.0-* family) |
per image, priced by model + resolution |
Bedrock Nova Canvas (amazon.nova-canvas-v1:0) |
per image, priced by quality (standard / premium) × width×height bucket |
Video generation
| Service |
Billable units |
| Azure OpenAI Sora |
per video-second at the requested size |
| Vertex Veo |
per video-second at the requested resolution (720p / 1080p / 4k) — audio adds a per-second up-charge on models that generate it |
| Bedrock Nova Reel |
per video-second at fixed 1280×720 / 24 fps |
Voice
| Service |
Billable units |
| Azure Speech (STT + TTS) |
STT: per hour of audio. TTS: per 1M characters (Neural), per 1M characters at a higher rate (HD / Multilingual / Dragon variants). |
| GCP Speech-to-Text v2 |
per minute of audio (billed per 15-sec increment); model tier changes rate |
| GCP Text-to-Speech |
per 1M characters (Standard) — different rate for Neural2, WaveNet, Chirp 3 HD |
| AWS Transcribe |
per second of audio (billed per second, 15-sec minimum), tier-dependent |
| AWS Polly |
per 1M characters — separate rates for standard, neural, long-form, generative engines |
Vector stores (from RAG)
| Service |
Billable units |
| Azure AI Search |
per replica-hour at the chosen SKU tier, plus per-document-storage-GB-month, plus vector-storage-GB-month on some SKUs |
| Vertex AI Vector Search |
per index-node-hour while an index is deployed, plus query cost per 1,000 queries, plus storage per-GB-month |
| Bedrock Knowledge Bases (backed by OpenSearch Serverless / Aurora / S3 Vectors) |
per underlying-store cost — OCU-hour for OpenSearch Serverless; ACU-hour for Aurora; per-GB-month + query cost for S3 Vectors |
| pgvector on managed Postgres |
per DB-instance-hour + storage-GB-month + network egress (no per-query charge) — see Store — pgvector |
| Service |
Billable units |
| AKS / GKE / EKS |
per node-hour (VM SKU rate), plus control-plane hourly on some tiers, plus egress |
| Azure Container Apps |
per vCPU-second + per GiB-second (memory), request-billed |
| Cloud Run |
per vCPU-second + per GiB-second, per request, per network egress |
| AWS App Runner |
per vCPU-hour + per GB-hour, per request |
The pattern
Every one of those rows is a stable metric. The rate shown at any URL above is not — refresh from the vendor's own page and calculator before committing a number to your own docs.