Skip to content

Native guardrails per cloud

The exact SDK call per cloud. All three surfaces return a verdict object (severity or block/allow) that the agnostic wrapper turns into Verdict(blocked=bool, reason=str).

Azure — AI Content Safety

Standalone screener (works for any model, not just Azure OpenAI). Send the prompt to analyze_text; the response has one severity per category (0, 2, 4, 6).

import os
from azure.ai.contentsafety import ContentSafetyClient
from azure.ai.contentsafety.models import AnalyzeTextOptions, TextCategory
from azure.core.credentials import AzureKeyCredential

client = ContentSafetyClient(
    endpoint=os.environ["CONTENT_SAFETY_ENDPOINT"],
    credential=AzureKeyCredential(os.environ["CONTENT_SAFETY_KEY"]),
)

response = client.analyze_text(AnalyzeTextOptions(
    text="Your input text",
    categories=[TextCategory.HATE, TextCategory.SELF_HARM,
                TextCategory.SEXUAL, TextCategory.VIOLENCE],
    blocklist_names=os.environ.get("CONTENT_SAFETY_BLOCKLISTS", "").split(",") or None,
    halt_on_blocklist_hit=True,
    output_type="FourSeverityLevels",     # or EightSeverityLevels
))

# Response shape: response.categories_analysis[i].severity in {0, 2, 4, 6}

For jailbreak / prompt-injection detection, use the Prompt Shields REST API (POST /contentsafety/text:shieldPrompt) — the response returns per-shield booleans (userPromptAnalysis.attackDetected, documents[i].attackDetected). The wrapper calls it in parallel with analyze_text on the input path.

Vertex — Gemini safety_settings (in-model)

The simplest guardrail is in-request — Gemini refuses to complete when a safety category exceeds threshold. Set thresholds per category on GenerateContentConfig:

from google import genai
from google.genai import types

client = genai.Client(vertexai=True, project=…, location=…)

resp = client.models.generate_content(
    model="gemini-3.6-flash",
    contents="Your prompt",
    config=types.GenerateContentConfig(
        safety_settings=[
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_HARASSMENT,
                threshold=types.HarmBlockThreshold.BLOCK_LOW_AND_ABOVE,
            ),
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
                threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
            ),
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT,
                threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
            ),
            types.SafetySetting(
                category=types.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
                threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
            ),
        ],
    ),
)

# If the model blocks: resp.candidates[0].finish_reason == FinishReason.SAFETY
# and resp.candidates[0].safety_ratings has the per-category verdicts.

Thresholds (from most-restrictive to least): BLOCK_LOW_AND_ABOVE, BLOCK_MEDIUM_AND_ABOVE, BLOCK_ONLY_HIGH, BLOCK_NONE, OFF.

Vertex — Model Armor (out-of-band policy layer)

For PII detection, prompt-injection defense, and denied topics beyond what Gemini's in-model filter does, use Model Armor — a separate policy-based screener (works with any LLM, not just Gemini):

from google.cloud import modelarmor_v1

client = modelarmor_v1.ModelArmorClient()

parent   = f"projects/{PROJECT}/locations/{LOCATION}"
template = f"{parent}/templates/{TEMPLATE_ID}"           # created out-of-band

# Screen an incoming prompt
req = modelarmor_v1.SanitizeUserPromptRequest(
    name=template,
    user_prompt_data=modelarmor_v1.DataItem(text="user prompt to screen"),
)
resp = client.sanitize_user_prompt(request=req)
# resp.sanitization_result.filter_match_state == MATCH_FOUND -> block

Screen the model output symmetrically with sanitize_model_response. Templates carry the enforcement mode (Inspect only vs Inspect and block) and per-category confidence thresholds (HIGH, MEDIUM_AND_ABOVE, LOW_AND_ABOVE).

Best practice from the docs: separate templates for prompts vs responses — the risk profiles differ.

AWS — Bedrock Guardrails

Two invocation styles:

Attached — the guardrail evaluates inline as part of converse/invoke_model:

resp = client.converse(
    modelId=…,
    messages=[…],
    guardrailConfig={
        "guardrailIdentifier": os.environ["BEDROCK_GUARDRAIL_ID"],
        "guardrailVersion":    os.environ.get("BEDROCK_GUARDRAIL_VERSION", "DRAFT"),
        "trace":               "enabled",
    },
)

Standalone — evaluate any text against the guardrail without invoking a model (useful for pre-model filtering or for wrapping non-Bedrock models):

import os, boto3

client = boto3.client("bedrock-runtime",
                      region_name=os.environ.get("AWS_REGION", "us-east-1"))

resp = client.apply_guardrail(
    guardrailIdentifier=os.environ["BEDROCK_GUARDRAIL_ID"],
    guardrailVersion=os.environ.get("BEDROCK_GUARDRAIL_VERSION", "DRAFT"),
    source="INPUT",                # or "OUTPUT" for model responses
    content=[{"text": {"text": "user prompt to screen"}}],
)

# resp["action"] == "GUARDRAIL_INTERVENED" -> block; assessments has the per-filter detail.

Policy shape (excerpt)

bedrock = boto3.client("bedrock", region_name=…)

bedrock.create_guardrail(
    name="chiron-guardrail",
    description="Chiron default guardrail",
    blockedInputMessaging="Sorry, that request violates policy.",
    blockedOutputsMessaging="Sorry, that response violates policy.",
    contentPolicyConfig={"filtersConfig": [
        {"type": "HATE",           "inputStrength": "HIGH", "outputStrength": "HIGH"},
        {"type": "INSULTS",        "inputStrength": "HIGH", "outputStrength": "HIGH"},
        {"type": "SEXUAL",         "inputStrength": "HIGH", "outputStrength": "HIGH"},
        {"type": "VIOLENCE",       "inputStrength": "HIGH", "outputStrength": "HIGH"},
        {"type": "MISCONDUCT",     "inputStrength": "HIGH", "outputStrength": "HIGH"},
        {"type": "PROMPT_ATTACK",  "inputStrength": "HIGH", "outputStrength": "NONE"},
    ]},
    topicPolicyConfig={"topicsConfig": [
        {"name": "financial_advice", "definition": "…", "type": "DENY"},
    ]},
    sensitiveInformationPolicyConfig={"piiEntitiesConfig": [
        {"type": "EMAIL",                       "action": "ANONYMIZE"},
        {"type": "PHONE",                       "action": "ANONYMIZE"},
        {"type": "US_SOCIAL_SECURITY_NUMBER",   "action": "BLOCK"},
    ]},
    contextualGroundingPolicyConfig={"filtersConfig": [
        {"type": "GROUNDING",  "threshold": 0.75},
        {"type": "RELEVANCE",  "threshold": 0.5},
    ]},
)

Common Verdict

The wrapper turns every native response into:

@dataclass
class Verdict:
    blocked: bool
    reason:  str        # human-readable, safe to surface to the caller
    raw:     dict       # native response (for logging via P2.6 observability)

The wrapper is the only place that touches native guardrail SDKs — the rest of the app treats guardrails as bool + str.