LLM guardrails
LLM guardrails: what they check, where they sit and where they fail
A plain account of what LLM guardrails are, the techniques behind them, their documented limits and the separate job of governance.
LLM guardrails are checks placed around a language model that decide whether an input, a retrieved passage, a tool call or an output should be blocked, changed or allowed. They reduce harm at request time. They are probabilistic, vendors say so in their own documentation, and they do not tell you who acted, under which policy or with what evidence.
Last verified 2026-10-07. Sources are listed at the end of the page.
What are LLM guardrails?
A guardrail is a check that runs between your application and a model, or between the model and the systems it touches. It answers one question: should this content be blocked, altered or passed on? NVIDIA describes its toolkit as a way to block, alter or validate unsafe, off-topic, malicious or policy-violating inputs and model responses. Amazon, Meta and Guardrails AI describe the same job with different mechanisms.
The word covers very different things. A regular expression that masks a card number is a guardrail. So is a trained classifier that labels a prompt as a hazard category, and so is a schema check that rejects a malformed JSON answer. They share a purpose but not a failure profile, which is why the table below separates them.
Where do guardrails sit in an LLM application?
NVIDIA lists five stages in its documentation, and most other products map onto a subset of them. The placement decides what the check can see.
Input
Runs before the model is called. It validates or sanitises user text. NVIDIA names content safety, jailbreak detection, topic control and PII masking as typical input uses.
Retrieval
Filters the documents or chunks a retrieval system returns, so only trusted context reaches the model. This is where indirect injection through a document can be caught, or missed.
Dialog flow
Constrains the shape of a multi-turn conversation, for example by keeping an assistant on its topic across turns. It needs state, so it is the hardest placement to test.
Tool call
Checks a tool or function call, its arguments and its result before the system acts. This is the placement closest to real-world harm, and the one that overlaps most with access control.
Output
Evaluates the model response before the user or a downstream system receives it. It can filter, edit or block, and it can validate structure as well as content.
Which guardrail technique catches what, and what does it miss?
| Guardrail type | Catches | Misses |
|---|---|---|
| Rules and regular expressions | Fixed shapes: key and token formats, card-like numbers, blocked words. Amazon Bedrock Guardrails word filters use case-insensitive whole-word and whole-phrase matching, and it also accepts custom regular expressions. | Paraphrase, spelling tricks, other languages and encoded text. OWASP lists encoded or multilingual instructions (for example Base64 or emojis) as a way to evade filters. |
| Trained classifiers | Content in the categories the model was trained on. Meta describes Llama Guard 4 as classifying both prompts and responses across 14 hazard categories, and Prompt Guard 2 as detecting prompt injection and jailbreak attempts. | Attacks outside the training distribution. Meta states that Llama Guard 4 may be susceptible to adversarial attacks or prompt injection, and warns that adversaries may develop attacks specifically to bypass Prompt Guard 2 detection. |
| LLM as judge | Context-dependent cases that rules cannot express, such as whether an answer breaches a written policy. | Anything that fools the judge, which is itself a model reading attacker-influenced text. Each judgement is also another model call, so it adds latency and cost. |
| Structured output validation | Malformed or out-of-range output. Guardrails AI lets you define the expected structure with Pydantic models and checks the response against it. | Harmful content that is well formed. A valid JSON field can still carry an injected instruction or a leaked secret. |
| Grounding and fact checks | Answers that drift from the retrieved source. Bedrock offers contextual grounding checks to flag responses not grounded in the source or irrelevant to the query. | Errors that the source itself contains, and attacks that arrive inside the source. |
| Tool call checks | Calls that break a rule on arguments, targets or scope, evaluated before execution. | Anything the rule did not anticipate. A check is only as good as the policy written into it, and the policy needs an owner. |
Descriptions of Bedrock Guardrails, Guardrails AI, Llama Guard 4 and Prompt Guard 2 are taken from each vendor page as read on 2026-10-07. This page does not compare their accuracy and publishes no detection rates.
How reliable are guardrails, and what do the vendors say?
Not reliable enough to be the only control. OWASP says it is unclear whether fool-proof prevention of prompt injection exists, because of the stochastic nature of generative models. NVIDIA writes in its security guidelines that its guardrails are not perfect, and that some safety rails can occasionally block otherwise safe requests, more often when several are combined.
Research agrees. Anthropic reported in February 2025 that its constitutional classifiers may not prevent every universal jailbreak, and that a public demonstration found one. Meta recommends Prompt Guard as an additional layer that complements other measures, and warns that adversaries may build attacks specifically to bypass detection. Read the jailbreak explainer for the attack families.
Two practical effects follow. First, a guardrail needs its own evaluation set, because you must measure false blocks as well as misses. Second, it must be treated as one layer. If a guardrail is the only thing between an injected instruction and a deleted record, the design is wrong, not the guardrail.
Guardrails or governance: what is the difference?
A guardrail asks whether something should be blocked. Governance asks who is acting, what they are allowed to do, under which policy, and with what evidence. You need both, and one does not replace the other.
| Question | Guardrail layer | Governance layer |
|---|---|---|
| Should this text be blocked? | Yes, this is its whole job. | Sets the policy the guardrail enforces, and reviews the results. |
| Who is acting? | Usually does not know. It sees a string. | Holds an identity for the person or agent, and an owner. |
| Allowed to do what? | Can refuse a tool call that breaks a rule it has been given. | Grants and limits permissions, and decides when a human must approve. |
| Under which policy? | Applies one rule set per application. | Versions policy, applies it per agent or per scope, and records which version ran. |
| With what evidence? | May log a block. | Keeps a record of the action, the decision and the approver that an auditor can read. |
How do you layer guardrails in practice?
1. Start from the action, not the text
List what the system can do if the model is fooled: send mail, write to a database, run a command. Limit those first. Guardrails help most where you cannot remove a capability.
2. Put cheap deterministic checks first
Use rules and schemas for what is fixed: secrets, identifiers, output structure. They are fast, repeatable and easy to explain to an auditor.
3. Add a classifier where the categories are clear
A hazard classifier suits content categories. Pick one whose model card matches your languages and domain, and follow the card on tuning for your own prompts.
4. Check tool calls before they run
Evaluate each proposed call against policy and run it under the permissions of the user, not of the model.
5. Measure false blocks and misses
Build a test set from your own traffic and from known attack patterns. Re-run it whenever you change a model, a prompt or a rule. The red-teaming guide covers the method.
6. Log decisions
Record what was blocked, what was allowed and why. See the AI audit trail page for what evidence is worth keeping.
Where Swfte fits
Connect, the model gateway, has a content-policy evaluator with built-in detectors for secrets and personal data, custom regular expressions, and a redact action that replaces a match with a placeholder (Built). Nexus, which wraps Claude Code and Codex, applies a policy gate with allow, deny and ask outcomes on tool calls, and records an audit trail (Built). The prompt injection page and agent runtime security describe how these sit with approvals.
The self-hosted model registry also lists a safety guard category, with Granite Guardian 3.3 8B among its entries. That is a catalogue entry for models you can host. It is not evidence of a detection pipeline, and Swfte publishes no detection rates and does not claim a jailbreak classifier.
You may not need Swfte for this. If your application calls one hosted model with no tools, the provider guardrails (for example Bedrock Guardrails, Llama Guard or NeMo Guardrails) plus your own output checks may be enough. Governance matters when agents act. See the AI governance page and the trust centre for what Swfte does and does not claim.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- NVIDIA NeMo Guardrails: rail types. The five stages and what each rail type does.
- NVIDIA NeMo Guardrails: security guidelines. Statements that guardrails are not perfect and can over-block.
- Guardrails AI on PyPI. Input and output guards, Hub validators and structured output.
- Llama Guard 4 model card. Prompt and response classification, hazard categories and stated limits.
- Llama Prompt Guard 2 model card. Injection and jailbreak detection and stated limits.
- Amazon Bedrock Guardrails. Filter types: content, denied topics, words, sensitive information, grounding.
- OWASP LLM01 Prompt Injection. Evasion by encoding and the statement on fool-proof prevention.
- Anthropic: constitutional classifiers. Classifier results and the stated limit on universal jailbreaks.
Frequently asked questions
What are LLM guardrails?
LLM guardrails are checks around a language model that block, change or allow inputs, retrieved content, tool calls and outputs. They can be rules, trained classifiers, a second model acting as judge, or schema validation. They reduce harm at request time but are probabilistic, so vendors recommend using them as one layer among several.
Are guardrails the same as prompt injection defence?
No. A guardrail may detect some injection attempts, but defence against injection also needs least privilege, approvals and egress limits, because OWASP says it is unclear whether fool-proof prevention exists. Detection lowers the rate of success. Limiting what a fooled model can do lowers the damage.
Can guardrails be bypassed?
Yes. Meta says Llama Guard 4 may be susceptible to adversarial attacks, and says adversaries may develop attacks specifically to bypass Prompt Guard. NVIDIA says its guardrails are not perfect. Anthropic reported that a public demonstration of its classifiers found one universal jailbreak. Test your own configuration and keep other controls in place.
Do guardrails replace human approval?
No. A guardrail can refuse a call that matches a rule, but it cannot judge intent in a new situation. OWASP lists human approval for high-risk operations as a separate mitigation. Use guardrails to filter and humans to decide on consequential or irreversible actions.
Which guardrail tool should I pick?
It depends on where you run and what you must catch. NeMo Guardrails and Guardrails AI are open-source Python libraries. Bedrock Guardrails is an AWS service. Llama Guard and Prompt Guard are models you host. Choose by language support, latency budget and the failure you fear most, then test it on your own data.
Does Swfte offer a jailbreak detector?
No claim is made. Swfte has a content-policy evaluator for secrets and personal data with a redact action, and a policy gate for coding agents. The self-hosted model registry lists safety guard models such as Granite Guardian 3.3 8B, but Swfte publishes no measured detection rates and does not claim jailbreak resistance.