AI jailbreak
AI jailbreak: what it is, the attack families and how to defend
A defensive explainer of AI jailbreaks: the definition, how they differ from prompt injection, the families researchers describe, and the layered defence that limits the damage.
An AI jailbreak is an attempt to make a model ignore its safety training or the policy it was given. OWASP treats it as a form of prompt injection, and researchers have shown it works against many models in several ways. No vendor claims a model is jailbreak-proof. The sound response is to limit what a jailbroken model can do and to watch for it. This page contains no attack prompts.
Last verified 2026-10-07. Sources are listed at the end of the page.
What is an AI jailbreak?
A jailbreak is input crafted to make a model act against its built-in safety behaviour. Meta's Prompt Guard 2 model card defines jailbreaks as malicious instructions designed to override the safety and security features built into a model. The target is the model's own training and policy.
The attacker is usually the person typing, who wants output the model was trained to refuse. A jailbreak becomes a security issue for a business when the model is connected to data or tools. A model that can be talked out of its policy can also be talked into misusing its access.
How is a jailbreak different from prompt injection?
OWASP states that jailbreaking is a form of prompt injection in which the attacker provides inputs that cause the model to disregard its safety protocols entirely. Meta lists the two as separate attack types. Both views are useful. The table separates them by where the attack aims.
| Question | Jailbreak | Prompt injection |
|---|---|---|
| What does it aim at? | The model's safety training and policy. | The application's instructions. Meta describes it as manipulating untrusted third-party and user data in the context window so the model runs unintended instructions. |
| Typical source | The user of the application. | The user, or content the application retrieves: a page, a document, a tool output. |
| Typical goal | Output the model would normally refuse. | A changed behaviour or action the application owner did not intend, such as data leaving the system. |
| Main defence | Model training, classifiers and output checks. | Least privilege, separating data from instructions, approval and egress control. |
See prompt injection defence for the second column in depth. Definitions are from OWASP LLM01 and the Meta Prompt Guard 2 model card, read on 2026-10-07.
Why do jailbreaks work at all?
Wei, Haghtalab and Steinhardt (2023) proposed two failure modes of safety training. Competing objectives arise when a model's capability and its safety goals pull against each other. Mismatched generalisation arises when safety training does not cover a domain where the model nonetheless has capability. They used these modes to design new attacks that succeeded against the models they tested, and they argue that safety mechanisms should be as sophisticated as the underlying model.
This explains why patching one phrase rarely ends the problem. The weakness is in how training generalises, so new phrasings keep appearing. The consequence for builders is plain: do not rely on the model alone to refuse.
Which jailbreak families do researchers describe?
Each family below is described at concept level from the cited research. None of these sources is needed to defend a system, and this page does not reproduce any working prompt.
Role-play and persona framing
The attacker asks the model to play a character that is not bound by its rules. Shen and colleagues (2023) analysed 1,405 jailbreak prompts collected from December 2022 to December 2023 and found one community built on turning the model into another character. They report that this community stopped sharing prompts after October 2023, possibly because vendors patched it.
Many-shot jailbreaking
The attacker fills a long context window with fabricated dialogues in which an assistant complies. Anthropic published this on 2024-04-02 and reported that the share of harmful responses rises as the number of shots rises, and that it was often more effective on larger models.
Multi-turn escalation
The attacker starts with a harmless question and escalates over several turns, referring to the model's own earlier replies. Russinovich, Salem and Eldan call this Crescendo and tested it on several public models. The paper was accepted at USENIX Security 2025.
Encoding and obfuscation
The attacker hides the request in a form safety training did not cover. Wei and colleagues explain that large models learn Base64 during pretraining, while safety training does not contain such unnatural input. OWASP lists multilingual and encoded instructions as a way to evade filters.
Adversarial suffixes
An automatically searched string of characters is appended to a request. Zou and colleagues (2023) produced such suffixes by combining greedy and gradient-based search, and reported that they transferred to public interfaces of ChatGPT, Bard and Claude and to several open-source models. OWASP lists the technique as an example scenario.
How do you defend against jailbreaks?
1. Assume it will sometimes work
Design for a model that has been talked out of its policy. List every action it could take and remove or gate the ones that matter. OWASP lists enforcing privilege control and least privilege access as a mitigation.
2. Screen inputs and outputs
Use classifiers and rules on both sides of the model. Anthropic reports that its constitutional classifier approach added false refusals and compute overhead, so budget for that. See LLM guardrails for the techniques.
3. Validate output formats and treat output as untrusted
OWASP recommends defining and validating expected output formats. Do not pass free-form output to a shell, a query or another system without a check.
4. Require approval for high-risk actions
A person approves sends, payments, deletions and publishing. See human in the loop.
5. Monitor and keep evidence
Log inputs, outputs, tool calls and policy decisions so a security team can investigate and tune. See the AI audit trail.
6. Red-team before launch and after every change
Test with current attack families, and re-test when you change a model, a prompt or a guardrail. Start with AI red teaming and the red-teaming how-to.
Is any model jailbreak-proof? What the vendors say
| Source | What it states (paraphrased) | Date |
|---|---|---|
| Anthropic, constitutional classifiers | Classifiers may not prevent every universal jailbreak. In a public demonstration four participants cleared all levels, and one found what Anthropic judged to be a universal jailbreak. Anthropic also reports a 0.38% rise in refusals and 23.7% extra compute cost. | Published 2025-02-03, updated to 2025-02-18 |
| Anthropic, many-shot jailbreaking | Fine-tuning to refuse merely delayed the jailbreak. Classifying and modifying the prompt before it reached the model cut one attack from 61% to 2%. Anthropic says it continues to develop defences. | Published 2024-04-02 |
| Meta, Prompt Guard 2 model card | Adversaries may develop attacks specifically to bypass detection. The card recommends the model as an additional layer that complements other measures. | Model card as read 2026-10-07 |
| Meta, Llama Guard 4 model card | As an LLM it may be susceptible to adversarial attacks or prompt injection that bypass or alter its intended use. | Model card as read 2026-10-07 |
| OWASP, LLM01 Prompt Injection | It is unclear whether fool-proof methods of prevention exist, because of the stochastic nature of generative models. | Page as read 2026-10-07 |
The figures are those the vendors published for their own tests and configurations. They are not a comparison of products and not a promise about your system.
Where Swfte fits
Swfte makes no claim that any model or agent running on the platform resists jailbreaks. The controls it does describe work on the consequences. Nexus applies a policy gate with allow, deny and ask outcomes to tool calls from Claude Code and Codex and records an audit trail (Built). Connect has a content-policy evaluator for secrets and personal data with a redact action (Built). The agent runtime security and AI red teaming pages describe the approach.
The Swfte Safety model card is design intent. It publishes no measured results, so it should not be read as evidence of jailbreak resistance. The card itself says a determined jailbreak remains possible.
You may not need Swfte if your system is a single chat feature with no tools and no private data. In that case the model provider's safeguards and your own output checks cover most of the exposure. Governance matters once the model can act. See the trust centre for what Swfte claims and does not.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- OWASP LLM01 Prompt Injection. Jailbreaking as a form of prompt injection, example scenarios and mitigations.
- Wei, Haghtalab and Steinhardt (2023): Jailbroken. Competing objectives and mismatched generalisation.
- Zou et al. (2023): Universal and Transferable Adversarial Attacks. Automatically produced adversarial suffixes and their transfer.
- Shen et al. (2023): "Do Anything Now". In-the-wild jailbreak prompts and persona-based communities.
- Russinovich, Salem and Eldan: Crescendo. Multi-turn escalation.
- Anthropic: many-shot jailbreaking. The technique, its scaling and the mitigation results.
- Anthropic: constitutional classifiers. Classifier defence, red-team results and stated limits.
- Llama Prompt Guard 2 model card. Definitions of jailbreak and injection, and stated limits.
- Llama Guard 4 model card. Statement on adversarial attacks against the classifier.
Frequently asked questions
What is an AI jailbreak?
An AI jailbreak is input designed to make a model ignore its built-in safety behaviour or the policy it was given. Meta defines jailbreaks as malicious instructions meant to override the safety and security features built into a model. OWASP classes jailbreaking as a form of prompt injection.
Is a jailbreak the same as prompt injection?
Not exactly. OWASP calls jailbreaking a form of prompt injection, while Meta treats them as two attack types. In practice a jailbreak targets the model's safety training, usually from the user, and prompt injection targets the application's instructions, often through retrieved content. The defences overlap but differ.
Can any model be made fully jailbreak-proof?
No vendor source read for this page claims so. Anthropic says its classifiers may not prevent every universal jailbreak, Meta warns that adversaries may build attacks to bypass its detection models, and OWASP says it is unclear whether fool-proof prevention exists. Plan for failure and limit what a compromised model can do.
What is many-shot jailbreaking?
It is a technique Anthropic described on 2024-04-02 in which a long prompt contains many fabricated dialogues of an assistant complying with requests, which raises the share of harmful responses as the number of examples grows. Anthropic reported that classifying and modifying the prompt cut one attack from 61% to 2%.
How do I test my own system?
Run a structured red-team exercise that covers the main families: persona framing, many-shot, multi-turn escalation, encoding and adversarial suffixes. Record what the model did, what the guardrails caught and what the connected tools allowed, then repeat after each model or prompt change. Use only systems you are authorised to test.
Does Swfte protect against jailbreaks?
Swfte makes no claim of jailbreak resistance. It offers controls that limit consequences, such as a policy gate on coding-agent tool calls and a content-policy evaluator for secrets and personal data. The Swfte Safety model card is design intent with no measured results, so it is not evidence of resistance.