Skip to content

Estimate — units × your current price

Chiron ships no dollar prices. This page teaches the arithmetic and the helper — you supply the current rates from Tooling.

Verified 2026-08-08

Every rate mentioned as an example below is a placeholder. Fetch real numbers from the pricing pages linked at the top of Units.

The formula

monthly_cost($) =  Σ (units_of_service_i  ×  current_rate_i)

That's it. The subtlety is (1) what the units are per service (see Units) and (2) getting fresh rates. Everything else is arithmetic.

Worked example — chat only

Suppose a Chiron call-an-llm chat service handles 10,000 requests/day, averaging 500 input tokens + 300 output tokens per turn. That's a stable estimate of units:

  • input tokens per month = 10,000 × 500 × 30 = 150,000,000 = 150 million
  • output tokens per month = 10,000 × 300 × 30 = 90,000,000 = 90 million

Multiply by whatever the current rate is on your chosen model, from the vendor's pricing page today. The estimator in examples/cost/service does the arithmetic; you supply the rates.

from estimator import estimate

report = estimate(
    units={
        "input_tokens_millions":  150,
        "output_tokens_millions":  90,
    },
    rates={
        # Fill from https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/
        "input_tokens_millions":  <YOUR_INPUT_RATE_USD_PER_1M>,
        "output_tokens_millions": <YOUR_OUTPUT_RATE_USD_PER_1M>,
    },
)
print(report)
# {
#   "total_usd": 150*<in_rate> + 90*<out_rate>,
#   "line_items": [
#     {"metric": "input_tokens_millions",  "units": 150, "rate": <in_rate>,  "cost_usd": ...},
#     {"metric": "output_tokens_millions", "units":  90, "rate": <out_rate>, "cost_usd": ...},
#   ],
# }

Worked example — RAG (chat + embed + pgvector)

Same 10,000 requests/day, plus embedding ingest at 100k docs / month (avg 250 input tokens/doc) and one small pgvector Postgres instance:

  • chat input: 150 M tokens/mo (as above)
  • chat output: 90 M tokens/mo
  • embed input: 10,000,000 tokens / 30 × 30 = 25 M tokens/mo (embed model rate is different — see units page)
  • pgvector db: 1 instance × 730 hours/mo × (its hourly rate)
  • pgvector storage: ~10 GB × 730 (hours) × (per-GB-hour rate) — most managed Postgres invoices per-GB-month

Feed those numbers to the estimator with your own rates. The service structure holds no matter which cloud you're on; only the rates and (sometimes) the units per-provider shift.

HTTP surface

The chart deploys the estimator as a stateless service:

POST /estimate
{
  "units": { "input_tokens_millions": 150, "output_tokens_millions": 90 },
  "rates": { "input_tokens_millions": 0.0, "output_tokens_millions": 0.0 }
}
→ 200
{
  "total_usd": 0.0,
  "line_items": [...],
  "note": "rates are supplied by the caller; Chiron does not carry vendor prices"
}

Wire it into your own dashboards, or call it from a build step to generate a per-PR "expected cost delta" comment.

Guardrail before you spend

Pair every estimate with a budget alert at 80% of the estimate. See Tooling — one Terraform resource per cloud. If the alert fires before month-end, either your units grew or the rate changed; investigate before the invoice.

The one thing to NOT do

Do not commit vendor rates to source. They will drift, they will get quoted back to you months later, and the estimate will be wrong when it matters. Fetch fresh from the pricing page every time.