Signals — what to record per LLM turn¶
Spans, metrics, and log events per turn. Attribute names follow the OpenTelemetry GenAI semantic conventions — pinning to that spec keeps every backend's charts and dashboards portable.
Verified 2026-08-08
- GenAI semantic conventions — opentelemetry.io/docs/specs/semconv/gen-ai
- OTel Python API — opentelemetry.io/docs/languages/python/instrumentation
Per-turn span¶
Wrap the model call in one span named after the operation. gen_ai.system identifies the provider; gen_ai.request.model and gen_ai.response.model are separate because the returned model can differ from the requested one (fallback / snapshot pinning).
from opentelemetry import trace
tracer = trace.get_tracer("chiron.observability")
with tracer.start_as_current_span("chat") as span:
span.set_attribute("gen_ai.system", "azure_openai") # azure_openai | vertex_ai | aws.bedrock
span.set_attribute("gen_ai.operation.name", "chat") # chat | text_completion | embeddings | image | video
span.set_attribute("gen_ai.request.model", "gpt-5")
span.set_attribute("gen_ai.request.max_tokens", 1024)
span.set_attribute("gen_ai.request.temperature", 0.7)
reply = azure_chat(prompt)
span.set_attribute("gen_ai.response.model", "gpt-5-2025-08-01")
span.set_attribute("gen_ai.usage.input_tokens", usage.input)
span.set_attribute("gen_ai.usage.output_tokens", usage.output)
# If your app has already looked up the current rate (see docs/cost/estimate.md),
# record the estimated cost as well. Otherwise omit — Chiron does not carry rates.
span.set_attribute("gen_ai.usage.cost_usd_estimate", cost_estimate)
Metrics¶
Three counters + one histogram cover most dashboards.
from opentelemetry import metrics
meter = metrics.get_meter("chiron.observability")
REQUESTS = meter.create_counter("gen_ai.client.requests",
unit="1", description="LLM calls issued")
INPUT_TOKENS = meter.create_counter("gen_ai.client.input_tokens", unit="{tokens}")
OUTPUT_TOKENS = meter.create_counter("gen_ai.client.output_tokens", unit="{tokens}")
LATENCY_MS = meter.create_histogram("gen_ai.client.duration", unit="ms")
# In your call site:
attrs = {"gen_ai.system": "azure_openai", "gen_ai.request.model": "gpt-5"}
REQUESTS.add(1, attrs)
INPUT_TOKENS.add(usage.input, attrs)
OUTPUT_TOKENS.add(usage.output, attrs)
LATENCY_MS.record(elapsed_ms, attrs)
Errors¶
An exception on the model call automatically records on the active span if you use with tracer.start_as_current_span(...) as span: — OTel calls span.record_exception(exc) and sets the status. To be explicit:
from opentelemetry.trace import Status, StatusCode
try:
reply = azure_chat(prompt)
except Exception as exc:
span.set_status(Status(StatusCode.ERROR, str(exc)))
span.record_exception(exc)
raise
Agent traces (multi-step)¶
Every tool call, retrieval step, and sub-LLM call becomes its own child span under the outer request span. OTel context propagation carries the parent — no extra wiring beyond start_as_current_span.
POST /rag ─┐
├─ retrieve (embed) ─── span: gen_ai.operation.name=embeddings
├─ retrieve (search) ── span: db.system=postgresql (pgvector)
└─ generate (chat) ─── span: gen_ai.operation.name=chat
├─ tool: get_current_temperature
└─ tool: lookup_docs
The E3 RAG rag_answer function already reads like a nested pipeline — wrap each of retrieve, _NATIVE[PROVIDER], and any tool invocation in its own span, and every backend renders the tree correctly.
Streaming¶
For a streaming chat, open the span at request-start, keep it open across chunks, and end it when the stream closes. time_to_first_token and time_to_last_token are useful histograms:
with tracer.start_as_current_span("chat.stream") as span:
t0 = time.monotonic()
first = True
for chunk in client.chat.completions.create(model=…, stream=True, messages=…):
if first:
span.set_attribute("gen_ai.response.time_to_first_token_ms",
int((time.monotonic() - t0) * 1000))
first = False
yield chunk
span.set_attribute("gen_ai.response.time_to_last_token_ms",
int((time.monotonic() - t0) * 1000))
Cost signal — the intentional gap¶
Chiron records gen_ai.usage.input_tokens + gen_ai.usage.output_tokens as first-class metrics, and only if your app already knows the rate also records gen_ai.usage.cost_usd_estimate. Rates come from Cost → Estimate, not from Chiron. Every dashboard that shows dollars should be a query that multiplies token metrics by your current rate at render-time — never a pre-baked price in the span.