Skip to content

Billable units per service

Prices drift; units don't. This table maps every service Chiron uses to the exact metric each cloud invoices.

Chat / embeddings / vision

Service Billable units
Azure OpenAI (GPT / o-series / embeddings / vision) per 1M input tokens, per 1M output tokens, per 1M cached input tokens (reduced), per-image tokens on multimodal (counted into the input-token bucket)
Vertex Gemini (text / image-in / audio-in / video-in) per 1M input tokens, per 1M output tokens, per 1M cached input tokens (reduced) — a different per-1M rate applies above the 200k-token boundary on some models
Bedrock (Anthropic / Nova / others) per 1M input tokens and per 1M output tokens (invoice line item: "on-demand invocations")

Vision input on all three is metered as tokens — the image is tokenized before the model reads it (Anthropic's rule is ⌈w/28⌉·⌈h/28⌉ visual tokens per image, capped by the model's tier — see Multimodal).

Image generation

Service Billable units
Azure OpenAI (gpt-image-*, historically DALL-E) per image, priced by size × quality × output_format
Vertex Imagen (imagen-4.0-* family) per image, priced by model + resolution
Bedrock Nova Canvas (amazon.nova-canvas-v1:0) per image, priced by quality (standard / premium) × width×height bucket

Video generation

Service Billable units
Azure OpenAI Sora per video-second at the requested size
Vertex Veo per video-second at the requested resolution (720p / 1080p / 4k) — audio adds a per-second up-charge on models that generate it
Bedrock Nova Reel per video-second at fixed 1280×720 / 24 fps

Voice

Service Billable units
Azure Speech (STT + TTS) STT: per hour of audio. TTS: per 1M characters (Neural), per 1M characters at a higher rate (HD / Multilingual / Dragon variants).
GCP Speech-to-Text v2 per minute of audio (billed per 15-sec increment); model tier changes rate
GCP Text-to-Speech per 1M characters (Standard) — different rate for Neural2, WaveNet, Chirp 3 HD
AWS Transcribe per second of audio (billed per second, 15-sec minimum), tier-dependent
AWS Polly per 1M characters — separate rates for standard, neural, long-form, generative engines

Vector stores (from RAG)

Service Billable units
Azure AI Search per replica-hour at the chosen SKU tier, plus per-document-storage-GB-month, plus vector-storage-GB-month on some SKUs
Vertex AI Vector Search per index-node-hour while an index is deployed, plus query cost per 1,000 queries, plus storage per-GB-month
Bedrock Knowledge Bases (backed by OpenSearch Serverless / Aurora / S3 Vectors) per underlying-store cost — OCU-hour for OpenSearch Serverless; ACU-hour for Aurora; per-GB-month + query cost for S3 Vectors
pgvector on managed Postgres per DB-instance-hour + storage-GB-month + network egress (no per-query charge) — see Store — pgvector

Compute (from Foundations)

Service Billable units
AKS / GKE / EKS per node-hour (VM SKU rate), plus control-plane hourly on some tiers, plus egress
Azure Container Apps per vCPU-second + per GiB-second (memory), request-billed
Cloud Run per vCPU-second + per GiB-second, per request, per network egress
AWS App Runner per vCPU-hour + per GB-hour, per request

The pattern

Every one of those rows is a stable metric. The rate shown at any URL above is not — refresh from the vendor's own page and calculator before committing a number to your own docs.