Agnostic wrapper — input → model → output¶
Small Python module that:
- Screens the input with the per-cloud native guardrail (fail closed on evaluator error).
- Calls the E2 chat service via
chat_native/chat_claudeimported verbatim (same pattern as Voice / RAG / Multimodal / Observability). - Screens the output with the per-cloud native guardrail (fail closed).
- Records a P2.6 OTel span per stage with the verdict — never the prompt/response body unless the redactor is proven active.
The load-bearing contract¶
class Verdict(NamedTuple):
blocked: bool
reason: str
raw: dict
class Guardrail(Protocol):
def screen_input(self, text: str) -> Verdict: ...
def screen_output(self, text: str) -> Verdict: ...
Three per-cloud implementations of Guardrail. Selection is CHIRON_PROVIDER — the same convention used everywhere else.
Fail-closed rule (Tony's hard rule)¶
try:
verdict = guardrail.screen_input(prompt)
except Exception as exc:
# Never treat evaluator failure as permission to bypass.
raise GuardrailError(f"input evaluator failed: {exc}") from exc
if verdict.blocked:
return {"blocked": True, "reason": verdict.reason}
The HTTP layer returns 502 on GuardrailError; the pod's readiness probe (from P2.6) already refuses to serve traffic if the guardrail client isn't loaded.
The chat function¶
def guarded_chat(prompt: str, *, stack: str = "native") -> dict:
guardrail = _dispatch_guardrail() # per-cloud
in_verdict = _try_screen(guardrail.screen_input, prompt, side="input")
if in_verdict.blocked:
return {"blocked": True, "side": "input", "reason": in_verdict.reason}
reply = (_NATIVE if stack == "native" else _CLAUDE)[PROVIDER](prompt)
out_verdict = _try_screen(guardrail.screen_output, reply, side="output")
if out_verdict.blocked:
return {"blocked": True, "side": "output", "reason": out_verdict.reason}
return {"blocked": False, "reply": reply}
_try_screen catches any exception from the evaluator and re-raises GuardrailError — the wrapper never silently drops to an "assume safe" branch.
Compose with observability¶
Every screen call is its own span; the verdict lands as an attribute so P2.6 dashboards can filter on guardrail.blocked == true:
with _tracer.start_as_current_span("guardrail.screen_input") as span:
span.set_attribute("guardrail.provider", PROVIDER)
span.set_attribute("guardrail.side", "input")
verdict = guardrail.screen_input(text)
span.set_attribute("guardrail.blocked", verdict.blocked)
span.set_attribute("guardrail.reason", verdict.reason)
# `raw` goes to logs only via the P2.6 redactor path — never as a span attr.
HTTP surface¶
POST /chat
{ "prompt": "...", "stack": "native" }
→
{ "blocked": true, "side": "input"|"output", "reason": "..." }
or
{ "blocked": false, "reply": "..." }
The 200 OK on a blocked response is deliberate — the wrapper is doing its job. If the caller wants HTTP-level distinction, wire the response's blocked flag to a 4xx in your gateway.
Attach the native guardrail into the model call (optional)¶
Bedrock and Azure OpenAI (via preview features) let you attach a guardrail to the model call itself. If you're already paying for a per-cloud native guardrail, using it inline is one fewer round-trip:
- Bedrock:
converse(..., guardrailConfig={"guardrailIdentifier": id, "guardrailVersion": v})— no separateapply_guardrailcall. - Azure OpenAI: content-filter deployments — set the content-filter policy on the deployment; requests inherit it.
Chiron's default is the sandwich pattern regardless, because it works for every model on every cloud, including through-line Claude — a portable defense-in-depth beats per-cloud shortcuts when you're teaching a pattern that has to survive a vendor migration.
The one thing to NOT do¶
Don't turn off screen_output. Model output is where PII from grounding data leaks, where jailbreaks succeed silently, and where hallucinated URLs sit. The input filter alone is not enough.