Skip to content

Signals — what to record per LLM turn

Spans, metrics, and log events per turn. Attribute names follow the OpenTelemetry GenAI semantic conventions — pinning to that spec keeps every backend's charts and dashboards portable.

Verified 2026-08-08

Per-turn span

Wrap the model call in one span named after the operation. gen_ai.system identifies the provider; gen_ai.request.model and gen_ai.response.model are separate because the returned model can differ from the requested one (fallback / snapshot pinning).

from opentelemetry import trace

tracer = trace.get_tracer("chiron.observability")

with tracer.start_as_current_span("chat") as span:
    span.set_attribute("gen_ai.system", "azure_openai")           # azure_openai | vertex_ai | aws.bedrock
    span.set_attribute("gen_ai.operation.name", "chat")           # chat | text_completion | embeddings | image | video
    span.set_attribute("gen_ai.request.model", "gpt-5")
    span.set_attribute("gen_ai.request.max_tokens", 1024)
    span.set_attribute("gen_ai.request.temperature", 0.7)

    reply = azure_chat(prompt)

    span.set_attribute("gen_ai.response.model", "gpt-5-2025-08-01")
    span.set_attribute("gen_ai.usage.input_tokens",  usage.input)
    span.set_attribute("gen_ai.usage.output_tokens", usage.output)
    # If your app has already looked up the current rate (see docs/cost/estimate.md),
    # record the estimated cost as well. Otherwise omit — Chiron does not carry rates.
    span.set_attribute("gen_ai.usage.cost_usd_estimate", cost_estimate)

Metrics

Three counters + one histogram cover most dashboards.

from opentelemetry import metrics

meter = metrics.get_meter("chiron.observability")

REQUESTS       = meter.create_counter("gen_ai.client.requests",
                                      unit="1", description="LLM calls issued")
INPUT_TOKENS   = meter.create_counter("gen_ai.client.input_tokens",  unit="{tokens}")
OUTPUT_TOKENS  = meter.create_counter("gen_ai.client.output_tokens", unit="{tokens}")
LATENCY_MS     = meter.create_histogram("gen_ai.client.duration",    unit="ms")

# In your call site:
attrs = {"gen_ai.system": "azure_openai", "gen_ai.request.model": "gpt-5"}
REQUESTS.add(1, attrs)
INPUT_TOKENS.add(usage.input,   attrs)
OUTPUT_TOKENS.add(usage.output, attrs)
LATENCY_MS.record(elapsed_ms,   attrs)

Errors

An exception on the model call automatically records on the active span if you use with tracer.start_as_current_span(...) as span: — OTel calls span.record_exception(exc) and sets the status. To be explicit:

from opentelemetry.trace import Status, StatusCode

try:
    reply = azure_chat(prompt)
except Exception as exc:
    span.set_status(Status(StatusCode.ERROR, str(exc)))
    span.record_exception(exc)
    raise

Agent traces (multi-step)

Every tool call, retrieval step, and sub-LLM call becomes its own child span under the outer request span. OTel context propagation carries the parent — no extra wiring beyond start_as_current_span.

POST /rag  ─┐
            ├─ retrieve (embed) ─── span: gen_ai.operation.name=embeddings
            ├─ retrieve (search) ── span: db.system=postgresql (pgvector)
            └─ generate (chat)  ─── span: gen_ai.operation.name=chat
                                    ├─ tool: get_current_temperature
                                    └─ tool: lookup_docs

The E3 RAG rag_answer function already reads like a nested pipeline — wrap each of retrieve, _NATIVE[PROVIDER], and any tool invocation in its own span, and every backend renders the tree correctly.

Streaming

For a streaming chat, open the span at request-start, keep it open across chunks, and end it when the stream closes. time_to_first_token and time_to_last_token are useful histograms:

with tracer.start_as_current_span("chat.stream") as span:
    t0 = time.monotonic()
    first = True
    for chunk in client.chat.completions.create(model=…, stream=True, messages=…):
        if first:
            span.set_attribute("gen_ai.response.time_to_first_token_ms",
                               int((time.monotonic() - t0) * 1000))
            first = False
        yield chunk
    span.set_attribute("gen_ai.response.time_to_last_token_ms",
                       int((time.monotonic() - t0) * 1000))

Cost signal — the intentional gap

Chiron records gen_ai.usage.input_tokens + gen_ai.usage.output_tokens as first-class metrics, and only if your app already knows the rate also records gen_ai.usage.cost_usd_estimate. Rates come from Cost → Estimate, not from Chiron. Every dashboard that shows dollars should be a query that multiplies token metrics by your current rate at render-time — never a pre-baked price in the span.