Skip to content

Content Safety & Guardrails

A prompt that produces the right format can still produce the wrong content, harmful, biased, or legally problematic output. Guardrails are safety layers that sit between the model and the user: they classify, filter, and rewrite outputs before anyone sees them. This lesson covers building guardrails that catch what prompt engineering alone can't prevent.

Gnome mascot

What you'll learn

  • Guardrails are output-side safety layers, they catch what the prompt can't prevent
  • Classification guardrails check for harmful categories before the output reaches the user
  • A layered approach catches more: keyword filter → classifier → human review for high-risk outputs

Build it

Step 1: Keyword-based guardrail (fastest, cheapest)

Keyword filters offer a lightweight first line of defense against straightforward attacks. They use hardcoded lists to rapidly block known bad terms.

python
BLOCKED_TERMS = {
    "pii": ["credit card", "social security", "ssn", "passport number"],
    "harmful": ["how to hack", "how to make explosives", "bypass security"],
    "legal": ["i am not a lawyer", "this is legal advice"],
}

def keyword_guardrail(output):
    for category, terms in BLOCKED_TERMS.items():
        for term in terms:
            if term in output.lower():
                return False, f"Blocked: {category} ({term})"
    return True, None

Step 2: LLM-as-classifier guardrail

When keyword lists fail to capture nuance, deploying a specialized language model as a classifier provides semantic analysis. This evaluates whether the generated output violates established safety policies before it reaches the user.

python
CLASSIFIER_PROMPT = """Classify the following AI output as SAFE or UNSAFE.

UNSAFE categories:
- HARMFUL: instructions for illegal activities, violence, self-harm
- HATE: discriminatory content, slurs, stereotyping
- PII: contains personal information like phone numbers, addresses, emails
- DECEPTIVE: impersonation, fraud, manipulation

Output: {output}

Respond with ONLY one word: SAFE or UNSAFE."""

def classifier_guardrail(output):
    result = call_llm(CLASSIFIER_PROMPT.format(output=output))
    return result.strip() == "SAFE"

Step 3: Layered guardrail pipeline

A robust production setup combines both approaches into a layered pipeline. This routes outputs through fast heuristic checks first, falling back to comprehensive model-based classification only when necessary.

python
def guardrail_pipeline(output):
    # Layer 1: Keyword (fast, cheap)
    passed, reason = keyword_guardrail(output)
    if not passed:
        return "[Response blocked by safety filter]", reason

    # Layer 2: Classifier (slower, more accurate)
    if not classifier_guardrail(output):
        return "[Response blocked by safety filter]", "classifier"

    # Layer 3: Human review flag (for high-risk categories)
    if is_high_risk(output):
        return "[Response flagged for human review]", "high_risk"

    return output, None

What goes wrong

MistakeHow you notice itThe fix
Keyword filter too aggressiveLegitimate outputs blocked ("credit" in financial context)Use regex with context, not substring matching. Add allowlists for known-safe phrases
Classifier slow for real-time useLatency doubles when guardrail is enabledRun keyword first (fast reject). Only invoke classifier on outputs that pass keyword but are in high-risk categories
Guardrail modifies output silentlyUsers confused why responses are differentLog every guardrail decision. Notify the user when output is modified: "This response has been modified by our safety system."

Confirm it worked

Execute the test cases below to ensure your guardrail functions as expected. You should verify that explicit attacks trigger the keyword block while safe statements pass through successfully.

python
# Test: harmful output blocked, safe output passes
harmful = "Here's how to hack into a bank account: first, you need to..."
assert not keyword_guardrail(harmful)[0]

safe = "The capital of France is Paris."
assert keyword_guardrail(safe)[0]

Next: Prompt Versioning & CI/CD