Observability & tracing¶
Instrument the E2 chat service + E3 RAG service once with OpenTelemetry, then export to whichever cloud you're on. The agnostic thread is OTel; the per-cloud backend is which exporter you attach.
The pipeline¶
┌────────────┐ ┌────────────────────┐ ┌──────────────────┐
│ your code │──►│ OpenTelemetry SDK │──►│ per-cloud │
│ + otel API │ │ (spans + metrics │ │ exporter or │
│ │ │ + logs, semantic │ │ collector │
│ │ │ conventions) │ │ │
└────────────┘ └────────────────────┘ └──────────────────┘
│
┌─────────────────────────┼─────────────────────────┐
▼ ▼ ▼
Azure Monitor Cloud Trace CloudWatch
(Application Insights) + Cloud Monitoring + X-Ray
Two layers per cloud¶
Each cloud offers two flavors of instrumentation — you almost always want both.
| Layer | What it tells you | Set up in |
|---|---|---|
| Application (OTel) | your app's own spans, prompt/response events, token counts, latency, errors | this recipe |
| Native model-invocation logging | every raw request/response to the LLM, captured server-side by the provider | one flag on the model resource — configured in the deploy modules |
Per-cloud OTel stacks¶
| Cloud | Python package(s) | Exporter class / setup one-liner |
|---|---|---|
| Azure | azure-monitor-opentelemetry |
configure_azure_monitor() — reads APPLICATIONINSIGHTS_CONNECTION_STRING env |
| GCP | opentelemetry-exporter-gcp-trace + opentelemetry-exporter-gcp-monitoring (in-process) or opentelemetry-exporter-otlp → OpenTelemetry Collector → Cloud Trace/Monitoring (recommended per GCP docs) |
Custom TracerProvider + CloudTraceSpanExporter, or OTLPSpanExporter → collector |
| AWS | aws-opentelemetry-distro (auto-instrumentation) |
OTEL_PYTHON_DISTRO=aws_distro OTEL_PYTHON_CONFIGURATOR=aws_configurator opentelemetry-instrument … (OTLP → ADOT Collector → X-Ray + CloudWatch) |
Native model-invocation logging per cloud¶
| Cloud | Native LLM audit trail |
|---|---|
| Azure OpenAI | Turn on diagnostic settings on the Cognitive Services account → route to a Log Analytics workspace (RequestResponse, Trace, AuditLogs categories). |
| Vertex AI | Enable Cloud Audit Logs for aiplatform.googleapis.com — Data Access logs capture predict/streamGenerateContent bodies (subject to redaction rules). |
| Bedrock | Turn on model invocation logging on the account: bedrock:PutModelInvocationLoggingConfiguration writes request/response JSON to CloudWatch Logs (or S3). |
Docs verified 2026-08-08¶
- OpenTelemetry Python — opentelemetry.io/docs/languages/python
- Azure Monitor OTel Distro — learn.microsoft.com/…/app/opentelemetry-enable
- GCP Cloud Trace + OTLP — cloud.google.com/trace/docs/setup/python-ot
- OTel GCP exporter package — pypi.org/project/opentelemetry-exporter-gcp-trace
- ADOT Python — aws-otel.github.io/docs/getting-started/python-sdk/auto-instr
- Semantic conventions for GenAI — opentelemetry.io/docs/specs/semconv/gen-ai
- Azure OpenAI diagnostic logs — learn.microsoft.com/…/openai/how-to/monitor-openai
- Vertex AI Cloud Audit Logs — cloud.google.com/vertex-ai/docs/general/audit-logging
- Bedrock model invocation logging — docs.aws.amazon.com/bedrock/latest/userguide/model-invocation-logging
Pages¶
- Signals — LLM-specific spans, metrics, events (tokens / latency / errors / agent-traces) and the GenAI semantic-conventions attribute names.
- Backends — per-cloud exporter setup — Application Insights / Cloud Trace + Monitoring / X-Ray + CloudWatch — plus native model-invocation logging.
- Prompt logging & privacy — when to log prompts + responses, what to redact first, and how each backend's log-scrubber lets you enforce it.
- Deploy — Terraform composing E1 baseline + the observability sink for each cloud (Log Analytics / Cloud Trace + Monitoring APIs / CloudWatch Log Group).
Validate-only¶
No live telemetry sent in CI. Each page's "verify" block is the local one-liner — run against your own creds to see the spans appear in the vendor console.