Skip to content

Guardrails & governance

The last Phase-2 topic. Wrap the E2 chat service with an input filter → model → output filter sandwich. The native guardrail per cloud does the heavy lifting; a small agnostic wrapper enforces the invariants that don't belong in any one vendor.

The pattern

     ┌──────────┐    ┌───────────────┐    ┌────────┐    ┌────────────────┐    ┌──────────┐
     │ prompt   │──► │ input filter  │──► │ E2     │──► │ output filter  │──► │  reply   │
     │ (text)   │    │ (per cloud +  │    │ chat   │    │ (per cloud +   │    │          │
     │          │    │  invariants)  │    │        │    │  invariants)   │    │          │
     └──────────┘    └───────────────┘    └────────┘    └────────────────┘    └──────────┘
                        │                                     │
                        └────── block on hit ─────────────────┘
                        └────── FAIL CLOSED on error ─────────┘

Hard rule: fail closed

If the guardrail service can't evaluate the prompt or the response, the wrapper blocks and returns 502. Never allow content through on evaluator error — an outage in the guardrail is not permission to bypass it.

Per-cloud native guardrails

Cloud Native surface What it covers
Azure Azure AI Content Safety — azure-ai-contentsafety ContentSafetyClient.analyze_text + shield_prompt for jailbreak/injection Hate / SelfHarm / Sexual / Violence with 4-level severity + custom blocklists + prompt-shields for injection
GCP Vertex Gemini safety_settings (in-model) + Model Armor (separate policy service, google-cloud-model-armor) In-model: 4 harm categories with 4 thresholds. Model Armor: harmful content + PII + prompt injection + denied topics + malicious URLs via templates
AWS Amazon Bedrock Guardrails — bedrock:CreateGuardrail + bedrock-runtime.apply_guardrail (standalone) or guardrailConfig in converse/invoke_model Content filters (Hate/Insults/Sexual/Violence/Misconduct/Prompt Attack) + denied topics + word filters + PII (block or mask) + contextual grounding for RAG

Docs verified 2026-08-08

Coverage matrix

Every column is a threat you want caught; every row is a native surface. - means "not directly covered — Chiron's agnostic layer handles it or you compose from other primitives."

Threat Azure Content Safety Vertex safety_settings + Model Armor Bedrock Guardrails
Moderation (hate / violence / sexual / self-harm) ✅ 4 cats × 4 levels ✅ 4 cats × 4 thresholds (Gemini native) ✅ Hate / Insults / Sexual / Violence / Misconduct
Custom blocklists / denied words ✅ blocklistNames ✅ (Model Armor) ✅ word filters
PII detection / redaction via Prompt Shields + custom regex ✅ Model Armor ✅ Sensitive info filter (block or mask)
Prompt-injection / jailbreak ✅ Prompt Shields (shield_prompt) ✅ Model Armor ✅ Prompt Attack filter
Denied topics (policy-level) via blocklists + custom classifier ✅ Model Armor templates ✅ denied topics
Contextual grounding (RAG) via app logic via app logic ✅ contextualGroundingPolicyConfig
Malicious URL screening via Prompt Shields ✅ Model Armor via app logic
Automated reasoning validation via app logic via app logic ✅ Automated Reasoning checks

Pages

  • Native guardrails — the per-cloud SDK calls (verified 2026-08-08).
  • Agnostic wrapper — the input → model → output sandwich around the E2 chat service. Fail-closed by design.
  • Deploy — Terraform composing E1 baseline + E2 call-an-llm + one guardrail resource per cloud.

Validate-only

No live guardrail evaluations in CI. The wrapper's /healthz reports which native surface is configured; unit tests run against a stub evaluator so the guardrail contract is exercised without money changing hands.