Guardrails & governance¶
The last Phase-2 topic. Wrap the E2 chat service with an input filter → model → output filter sandwich. The native guardrail per cloud does the heavy lifting; a small agnostic wrapper enforces the invariants that don't belong in any one vendor.
The pattern¶
┌──────────┐ ┌───────────────┐ ┌────────┐ ┌────────────────┐ ┌──────────┐
│ prompt │──► │ input filter │──► │ E2 │──► │ output filter │──► │ reply │
│ (text) │ │ (per cloud + │ │ chat │ │ (per cloud + │ │ │
│ │ │ invariants) │ │ │ │ invariants) │ │ │
└──────────┘ └───────────────┘ └────────┘ └────────────────┘ └──────────┘
│ │
└────── block on hit ─────────────────┘
└────── FAIL CLOSED on error ─────────┘
Hard rule: fail closed¶
If the guardrail service can't evaluate the prompt or the response, the wrapper blocks and returns 502. Never allow content through on evaluator error — an outage in the guardrail is not permission to bypass it.
Per-cloud native guardrails¶
| Cloud | Native surface | What it covers |
|---|---|---|
| Azure | Azure AI Content Safety — azure-ai-contentsafety ContentSafetyClient.analyze_text + shield_prompt for jailbreak/injection |
Hate / SelfHarm / Sexual / Violence with 4-level severity + custom blocklists + prompt-shields for injection |
| GCP | Vertex Gemini safety_settings (in-model) + Model Armor (separate policy service, google-cloud-model-armor) |
In-model: 4 harm categories with 4 thresholds. Model Armor: harmful content + PII + prompt injection + denied topics + malicious URLs via templates |
| AWS | Amazon Bedrock Guardrails — bedrock:CreateGuardrail + bedrock-runtime.apply_guardrail (standalone) or guardrailConfig in converse/invoke_model |
Content filters (Hate/Insults/Sexual/Violence/Misconduct/Prompt Attack) + denied topics + word filters + PII (block or mask) + contextual grounding for RAG |
Docs verified 2026-08-08¶
- Azure AI Content Safety Python — learn.microsoft.com/…/content-safety/quickstart-text
- Azure Prompt Shields (jailbreak) — learn.microsoft.com/…/content-safety/quickstart-jailbreak
- Vertex Gemini safety settings — ai.google.dev/gemini-api/docs/safety-settings
- Google Cloud Model Armor overview — cloud.google.com/security-command-center/docs/model-armor-overview
- Bedrock Guardrails — docs.aws.amazon.com/bedrock/latest/userguide/guardrails
- Bedrock ApplyGuardrail API — docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ApplyGuardrail
Coverage matrix¶
Every column is a threat you want caught; every row is a native surface. - means "not directly covered — Chiron's agnostic layer handles it or you compose from other primitives."
| Threat | Azure Content Safety | Vertex safety_settings + Model Armor |
Bedrock Guardrails |
|---|---|---|---|
| Moderation (hate / violence / sexual / self-harm) | ✅ 4 cats × 4 levels | ✅ 4 cats × 4 thresholds (Gemini native) | ✅ Hate / Insults / Sexual / Violence / Misconduct |
| Custom blocklists / denied words | ✅ blocklistNames | ✅ (Model Armor) | ✅ word filters |
| PII detection / redaction | via Prompt Shields + custom regex | ✅ Model Armor | ✅ Sensitive info filter (block or mask) |
| Prompt-injection / jailbreak | ✅ Prompt Shields (shield_prompt) | ✅ Model Armor | ✅ Prompt Attack filter |
| Denied topics (policy-level) | via blocklists + custom classifier | ✅ Model Armor templates | ✅ denied topics |
| Contextual grounding (RAG) | via app logic | via app logic | ✅ contextualGroundingPolicyConfig |
| Malicious URL screening | via Prompt Shields | ✅ Model Armor | via app logic |
| Automated reasoning validation | via app logic | via app logic | ✅ Automated Reasoning checks |
Pages¶
- Native guardrails — the per-cloud SDK calls (verified 2026-08-08).
- Agnostic wrapper — the
input → model → outputsandwich around the E2 chat service. Fail-closed by design. - Deploy — Terraform composing E1 baseline + E2 call-an-llm + one guardrail resource per cloud.
Validate-only¶
No live guardrail evaluations in CI. The wrapper's /healthz reports which native surface is configured; unit tests run against a stub evaluator so the guardrail contract is exercised without money changing hands.