Skip to content

Agnostic wrapper — input → model → output

Small Python module that:

  1. Screens the input with the per-cloud native guardrail (fail closed on evaluator error).
  2. Calls the E2 chat service via chat_native / chat_claude imported verbatim (same pattern as Voice / RAG / Multimodal / Observability).
  3. Screens the output with the per-cloud native guardrail (fail closed).
  4. Records a P2.6 OTel span per stage with the verdict — never the prompt/response body unless the redactor is proven active.

The load-bearing contract

class Verdict(NamedTuple):
    blocked: bool
    reason:  str
    raw:     dict


class Guardrail(Protocol):
    def screen_input(self,  text: str) -> Verdict: ...
    def screen_output(self, text: str) -> Verdict: ...

Three per-cloud implementations of Guardrail. Selection is CHIRON_PROVIDER — the same convention used everywhere else.

Fail-closed rule (Tony's hard rule)

try:
    verdict = guardrail.screen_input(prompt)
except Exception as exc:
    # Never treat evaluator failure as permission to bypass.
    raise GuardrailError(f"input evaluator failed: {exc}") from exc

if verdict.blocked:
    return {"blocked": True, "reason": verdict.reason}

The HTTP layer returns 502 on GuardrailError; the pod's readiness probe (from P2.6) already refuses to serve traffic if the guardrail client isn't loaded.

The chat function

def guarded_chat(prompt: str, *, stack: str = "native") -> dict:
    guardrail = _dispatch_guardrail()             # per-cloud

    in_verdict = _try_screen(guardrail.screen_input, prompt, side="input")
    if in_verdict.blocked:
        return {"blocked": True, "side": "input", "reason": in_verdict.reason}

    reply = (_NATIVE if stack == "native" else _CLAUDE)[PROVIDER](prompt)

    out_verdict = _try_screen(guardrail.screen_output, reply, side="output")
    if out_verdict.blocked:
        return {"blocked": True, "side": "output", "reason": out_verdict.reason}

    return {"blocked": False, "reply": reply}

_try_screen catches any exception from the evaluator and re-raises GuardrailError — the wrapper never silently drops to an "assume safe" branch.

Compose with observability

Every screen call is its own span; the verdict lands as an attribute so P2.6 dashboards can filter on guardrail.blocked == true:

with _tracer.start_as_current_span("guardrail.screen_input") as span:
    span.set_attribute("guardrail.provider", PROVIDER)
    span.set_attribute("guardrail.side",     "input")
    verdict = guardrail.screen_input(text)
    span.set_attribute("guardrail.blocked",  verdict.blocked)
    span.set_attribute("guardrail.reason",   verdict.reason)
    # `raw` goes to logs only via the P2.6 redactor path — never as a span attr.

HTTP surface

POST /chat
{ "prompt": "...", "stack": "native" }
→
  { "blocked": true, "side": "input"|"output", "reason": "..." }
  or
  { "blocked": false, "reply": "..." }

The 200 OK on a blocked response is deliberate — the wrapper is doing its job. If the caller wants HTTP-level distinction, wire the response's blocked flag to a 4xx in your gateway.

Attach the native guardrail into the model call (optional)

Bedrock and Azure OpenAI (via preview features) let you attach a guardrail to the model call itself. If you're already paying for a per-cloud native guardrail, using it inline is one fewer round-trip:

  • Bedrock: converse(..., guardrailConfig={"guardrailIdentifier": id, "guardrailVersion": v}) — no separate apply_guardrail call.
  • Azure OpenAI: content-filter deployments — set the content-filter policy on the deployment; requests inherit it.

Chiron's default is the sandwich pattern regardless, because it works for every model on every cloud, including through-line Claude — a portable defense-in-depth beats per-cloud shortcuts when you're teaching a pattern that has to survive a vendor migration.

The one thing to NOT do

Don't turn off screen_output. Model output is where PII from grounding data leaks, where jailbreaks succeed silently, and where hallucinated URLs sit. The input filter alone is not enough.