Skip to content

Observability & tracing

Instrument the E2 chat service + E3 RAG service once with OpenTelemetry, then export to whichever cloud you're on. The agnostic thread is OTel; the per-cloud backend is which exporter you attach.

The pipeline

     ┌────────────┐   ┌────────────────────┐   ┌──────────────────┐
     │ your code  │──►│  OpenTelemetry SDK │──►│  per-cloud       │
     │ + otel API │   │  (spans + metrics  │   │  exporter or     │
     │            │   │   + logs, semantic │   │  collector       │
     │            │   │   conventions)     │   │                  │
     └────────────┘   └────────────────────┘   └──────────────────┘
                                                        │
                              ┌─────────────────────────┼─────────────────────────┐
                              ▼                         ▼                         ▼
                       Azure Monitor              Cloud Trace              CloudWatch
                     (Application Insights)     + Cloud Monitoring        + X-Ray

Two layers per cloud

Each cloud offers two flavors of instrumentation — you almost always want both.

Layer What it tells you Set up in
Application (OTel) your app's own spans, prompt/response events, token counts, latency, errors this recipe
Native model-invocation logging every raw request/response to the LLM, captured server-side by the provider one flag on the model resource — configured in the deploy modules

Per-cloud OTel stacks

Cloud Python package(s) Exporter class / setup one-liner
Azure azure-monitor-opentelemetry configure_azure_monitor() — reads APPLICATIONINSIGHTS_CONNECTION_STRING env
GCP opentelemetry-exporter-gcp-trace + opentelemetry-exporter-gcp-monitoring (in-process) or opentelemetry-exporter-otlp → OpenTelemetry Collector → Cloud Trace/Monitoring (recommended per GCP docs) Custom TracerProvider + CloudTraceSpanExporter, or OTLPSpanExporter → collector
AWS aws-opentelemetry-distro (auto-instrumentation) OTEL_PYTHON_DISTRO=aws_distro OTEL_PYTHON_CONFIGURATOR=aws_configurator opentelemetry-instrument … (OTLP → ADOT Collector → X-Ray + CloudWatch)

Native model-invocation logging per cloud

Cloud Native LLM audit trail
Azure OpenAI Turn on diagnostic settings on the Cognitive Services account → route to a Log Analytics workspace (RequestResponse, Trace, AuditLogs categories).
Vertex AI Enable Cloud Audit Logs for aiplatform.googleapis.com — Data Access logs capture predict/streamGenerateContent bodies (subject to redaction rules).
Bedrock Turn on model invocation logging on the account: bedrock:PutModelInvocationLoggingConfiguration writes request/response JSON to CloudWatch Logs (or S3).

Docs verified 2026-08-08

Pages

  • Signals — LLM-specific spans, metrics, events (tokens / latency / errors / agent-traces) and the GenAI semantic-conventions attribute names.
  • Backends — per-cloud exporter setup — Application Insights / Cloud Trace + Monitoring / X-Ray + CloudWatch — plus native model-invocation logging.
  • Prompt logging & privacy — when to log prompts + responses, what to redact first, and how each backend's log-scrubber lets you enforce it.
  • Deploy — Terraform composing E1 baseline + the observability sink for each cloud (Log Analytics / Cloud Trace + Monitoring APIs / CloudWatch Log Group).

Validate-only

No live telemetry sent in CI. Each page's "verify" block is the local one-liner — run against your own creds to see the spans appear in the vendor console.