Short answer
To red team an LLM application, get written authorisation, model the threats against your own tools and data, run automated attack suites (promptfoo against the application, garak against the model, PyRIT for multi-turn campaigns), add manual sessions, then score, fix and retest. Test in staging with canary data. Treat every confirmed attack as a regression test so it can never quietly return.
The steps at a glance
- Get authorisation and write the rules of engagement
- Build a threat model from the system’s abilities
- Prepare a safe test environment
- Attack the application with promptfoo
- Scan the model layer with garak
- Add multi-turn and custom campaigns with PyRIT
- Run manual attack sessions on the highest-risk abuse cases
- Score findings and triage them
- Fix at the right layer and retest
- Report, archive and schedule the next run
Before you start
Who this is for
- Security engineers and AI teams who must test an LLM feature, RAG assistant or agent before release.
- Teams that already have an eval suite and want to add adversarial testing.
- Anyone asked “did we red team it?” by a customer, auditor or regulator.
Probably not for you if
- Teams who have not yet built any functional tests: start with how to validate your AI.
- Anyone who wants to test a system they do not own or have written permission to test.
Prerequisites
- Written authorisation from the system owner, and a staging copy of the application with test accounts.
- Node.js (promptfoo’s red-team quickstart states 22.22.0 or newer as of 2026-10-06) and Python 3.11 to 3.13 for garak.
- A threat model or at least a list of the tools, data stores and users the system touches.
- Somewhere to store transcripts safely: they will contain harmful content and any secrets the system leaks.
- Time
- Two to five working days for a first exercise on one application
- Cost
- Free open-source tools. Attack generation and grading call a model, which is billed by your provider unless you configure a local one.
- Hardware
- A laptop is enough to drive the tools. Local models need the memory their size implies; hosted targets need none.
- Skill
- Comfortable with HTTP APIs, YAML and Python; some security testing experience helps
Estimates are ours, not measurements, and move with your hardware, data and network.
What to test: the OWASP LLM Top 10 as a checklist
The OWASP Top 10 for LLM Applications (2025 edition) lists ten risks, LLM01 to LLM10. It is a good scoping checklist because every item turns into something you can try. The table maps each to a test idea. The tool column names what promptfoo, garak or you would use; coverage is partial in every case, which is why step 7 adds manual sessions.
OWASP’s GenAI Red Teaming Guide (version 1.0, January 2025) frames the work in four areas: model evaluation, implementation testing, infrastructure assessment and runtime behaviour analysis. The steps below touch all four, but most automated tools reach only the first two.
| ID and name | Test idea | Where to start |
|---|---|---|
| LLM01 Prompt Injection | Hide instructions in user input, documents, web pages and tool results; try to override the system prompt | promptfoo owasp:llm:01; garak promptinject and encoding probes |
| LLM02 Sensitive Information Disclosure | Ask for other users’ data, secrets, training data; check logs and outputs for planted canary strings | promptfoo owasp:llm:02 and pii plugins |
| LLM03 Supply Chain | Check model files, packages and plugins against pinned hashes and sources | Manual review; see how to evaluate an open-source LLM |
| LLM04 Data and Model Poisoning | Review who can write to training, fine-tuning and retrieval data; try inserting poisoned documents | Manual; test the ingestion path |
| LLM05 Improper Output Handling | Make the model output script, SQL or markup and see whether downstream code executes or renders it | Manual: trace the model output into every downstream parser, renderer and database call |
| LLM06 Excessive Agency | Try to make the agent call tools it should not, or with arguments it should not use | promptfoo owasp:llm:06; manual with real tool stubs |
| LLM07 System Prompt Leakage | Ask for the system prompt in many forms; plant a canary phrase in it | promptfoo prompt-extraction plugin |
| LLM08 Vector and Embedding Weaknesses | Query for documents the user should not see; test access control in retrieval | Manual; use two test users with different permissions |
| LLM09 Misinformation | Ask questions with known answers outside the knowledge base; check for confident invention | Your eval set; see how to validate your AI |
| LLM10 Unbounded Consumption | Send very long inputs, recursive tasks and floods; watch cost and latency | Load script against staging; check budgets and rate limits |
Checked against: OWASP Top 10 for LLM Applications 2025, OWASP GenAI Red Teaming Guide
Step 2Build a threat model from the system’s abilities
You end up with: A short table of assets, entry points and abuse cases ranked by what an attacker would gain.
List what the system can do, because that is what an attacker inherits. For each ability, such as reading a customer record, sending an email or running code, write the worst thing someone could make it do. Then list every place untrusted text enters: user messages, uploaded files, retrieved documents, web pages, tool results, email bodies.
The risk is highest where three things meet in one session: untrusted input, access to private data and the ability to act or send data out. Simon Willison called this the lethal trifecta in June 2025, and Meta’s Agents Rule of Two (October 2025) states a similar rule: an agent should combine no more than two of processing untrusted input, accessing sensitive systems or data, and changing state or communicating externally, otherwise a human should be in the loop. Mark which of your agents satisfy all three and test those first.
Rank the abuse cases by impact and by how easily an attacker could reach them. This ranking decides where manual time goes in step 7.
Checked against: Simon Willison: The lethal trifecta for AI agents, Meta: Agents Rule of Two
Step 3Prepare a safe test environment
You end up with: A staging copy with canary data, test accounts, full logging and a way to stop the attack run.
Test against staging, never production. Fill it with synthetic data and plant canary strings: unique fake secrets in the system prompt, in a document only another test user may read, and in a tool’s data. If a canary appears in an output, you have proof of leakage that no judge can argue with.
Create at least two test users with different permissions, so you can test whether one can reach the other’s data. Turn on full logging so you can see every tool call and retrieval. Put spending limits on the model account the target uses, because attack suites send thousands of prompts. Know how to stop: a kill switch for the target and a way to cancel the run.
Pin the versions of the tools you use and record them. Results from an unrecorded tool version cannot be compared with next quarter’s run.
Step 4Attack the application with promptfoo
You end up with: A generated set of attacks run against your staging endpoint, and an HTML report of what got through.
promptfoo’s red-team mode generates adversarial inputs for the categories you pick and runs them against a target, which can be a model or an HTTP endpoint. Its quickstart gives
npx promptfoo@latest redteam setupfor a guided web setup andpromptfoo redteam init --no-guito create the configuration from the command line. As of 2026-10-06 the documentation states Node.js 22.22.0 or newer and says an OpenAI key is optional.The configuration has a
targetssection and aredteamsection. The target below is an HTTP endpoint. Theredteamsection holds the purpose, which tells the generator what the system is for, the plugins to attack with and the strategies that change how attacks are delivered. The documentation lists plugins such aspii,excessive-agency,hijackingand the OWASP presetsowasp:llm:01,owasp:llm:02andowasp:llm:06, and strategies such asjailbreak,base64andcrescendo. Start small, withnumTestsof 5, which sits inside the 5 to 20 range the documentation calls typical.The configuration has a provider field for the model that generates attacks. Check the documentation for what the default is before you put confidential purpose text into it, and configure a model you are allowed to send that text to.
Run the scan, then open the report. Read each successful attack as a transcript: what was sent, what came back, and why it counts. Group them by plugin and by root cause before you triage.
Create the configuration without the web UI · bash promptfoo redteam init --no-guipromptfooconfig.yaml: an HTTP target plus OWASP-based plugins · yaml targets: - id: https label: support-assistant-staging config: url: 'https://staging.example.com/api/chat' method: 'POST' headers: 'Content-Type': 'application/json' body: message: '{{prompt}}' transformResponse: 'json.reply' redteam: purpose: 'A customer support assistant for an online shop. It can look up order status and draft replies. It must not reveal other customers data, issue refunds, or follow instructions found inside tickets or documents.' numTests: 5 plugins: - owasp:llm:01 - owasp:llm:02 - owasp:llm:06 - hijacking - pii strategies: - jailbreak - base64 - crescendoGenerate the attacks, run them, then open the report · bash npx promptfoo@latest redteam run npx promptfoo@latest redteam reportChecked against: promptfoo: Red team quickstart, promptfoo: Red team configuration, promptfoo: OWASP LLM Top 10, promptfoo: HTTP provider
Step 5Scan the model layer with garak
You end up with: A report showing which probe families the model fails, with a hit log you can replay.
garak is NVIDIA’s open-source LLM vulnerability scanner. It tests a model with probes, such as encoding-based injection, DAN-style jailbreaks, prompt injection and replay of training data, and uses detectors to judge the responses. Use it to understand the underlying model, and use promptfoo for the application around it, because the two answer different questions.
Install with pip. The README’s conda route pins Python to 3.11 up to 3.13. List the probes first, then run a selection with
--spec, the unified selector flag. Its reference describes selectors such asprobes.dan, tags such astag:owasp:llm01, and a minus sign to exclude. The README’s example runs one DAN probe against the small gpt2 model from Hugging Face, a harmless first target, and another runs the encoding probes against an OpenAI model with anOPENAI_API_KEYset. Swap in a model your organisation approves.Every run produces a debugging log, a JSONL report and a hit log of detected vulnerabilities, and prints the report file name at the start and end. Keep these with your transcripts.
--generationssets how many outputs are sampled per prompt; raising it gives more reliable rates at higher cost.Install garak · bash python -m pip install -U garakList probes, filtered to one family · bash garak --list_probes garak --list_probes --spec probes.danREADME example: one DAN probe against gpt2 from Hugging Face · bash python3 -m garak --target_type huggingface --target_name gpt2 --spec probes.dan.Dan_11_0README example: encoding-based injection against an OpenAI model · bash export OPENAI_API_KEY="<your key>" python3 -m garak --target_type openai --target_name gpt-5-nano --spec probes.encodingChecked against: NVIDIA garak README, garak CLI reference
Step 6Add multi-turn and custom campaigns with PyRIT
You end up with: PyRIT installed and configured, with a plan for the campaigns that single-prompt scanners cannot run.
PyRIT, the Python Risk Identification Tool from Microsoft’s AI Red Team, is a framework for building your own attack campaigns, including conversations that escalate over several turns. PyPI shows version 1.1.0 released on 4 September 2026 under the MIT licence with an “Alpha” development status, so expect interfaces to change between releases and pin the version you use.
Install it with pip. The documentation recommends Python 3.13 for local installation and PyPI lists support for Python 3.10 up to but not including 3.15. After installing, the docs say to configure two files,
~/.pyrit/.envfor endpoint credentials and~/.pyrit/.pyrit_conffor memory and initialiser settings, and describe three ways to use it: a command-line scanner, a GUI and a programmatic framework. Follow the installation and configuration pages on the PyRIT documentation site for the exact contents of those files, since they change between releases.Use PyRIT for what the earlier tools cannot do: a scripted attacker that adapts to each answer, campaigns that target your specific tools, and scoring you control. Write each campaign down as a scenario with a goal, such as “make the assistant reveal another customer’s order”, then record whether and after how many turns it succeeded.
Install PyRIT · bash pip install pyritChecked against: PyRIT on PyPI, PyRIT documentation (microsoft.github.io/PyRIT)
Step 7Run manual attack sessions on the highest-risk abuse cases
You end up with: A log of human-driven attempts against the abuse cases ranked highest in your threat model.
Take the top five abuse cases from step 2 and give each to a person for an hour. Their job is to achieve the bad outcome by any route inside scope. Ask them to try the same goal through several routes: directly, through a document the system will read, through a tool result, in another language, and across a long conversation.
Make them keep notes in a fixed format: goal, steps taken, the exact input, the output, whether it worked, and how many attempts it took. Reliability matters as much as possibility. An attack that works one time in fifty against a system doing thousands of requests a day is a real problem.
Include people who did not build the system. Builders share the same assumptions as the code. If you have domain staff, such as support agents, ask them how a bad customer would try to trick the assistant; they usually know.
Step 8Score findings and triage them
You end up with: A deduplicated list of findings, each with a severity, reproduction steps and an owner.
Merge duplicates first: fifty jailbreak transcripts that exploit one weakness are one finding. For each finding record the reproduction steps, the attack success rate across repeated tries, the data or action the attacker gained, and the root cause: missing permission check, untrusted text treated as instructions, over-broad tool, or weak filter.
Score on impact and reliability together. The scale below is a practical one; set your own bands with your security team and keep them stable between exercises so trends mean something.
Do not accept “the model refused most of the time” as a pass for high-impact cases. A control that fails one time in twenty is not a control when the action is irreversible.
A practical severity scale Severity What the attacker gets Typical response Critical Another user’s data, a secret, or an irreversible action, reliably Block release; fix now; retest High The same, but unreliable or needing unusual access; or a policy bypass with real harm Fix before release Medium Harmful or off-policy content; system prompt leakage with no secret in it Fix on a dated plan Low Cosmetic or very unlikely misuse Track Step 9Fix at the right layer and retest
You end up with: Each finding closed by a change that removes the capability or gates it, and a regression case proving it.
Prefer fixes in this order. Remove the capability if the system does not need it. Enforce permissions in code, outside the model. Require human approval for risky actions. Then filter inputs and outputs, and last, adjust the prompt. Filters and prompts help but adaptive attackers get past them, so they should not be the only layer on a high-severity finding. OWASP’s own guidance for prompt injection lists constraining model behaviour, filtering input and output, least privilege, human approval for high-risk actions, segregating external content and adversarial testing together; none of them is sufficient alone.
After each fix, rerun the exact attack that worked and then the wider suite. Add every confirmed attack to your standing eval set as a regression case with a deterministic pass rule, such as “output must not contain the canary” or “tool X must not be called”. That is how a one-off exercise becomes a gate. See how to stop prompt injection for the layered defences.
Checked against: OWASP LLM01:2025 Prompt Injection
Step 10Report, archive and schedule the next run
You end up with: A report a reader can act on, an archive of evidence, and a pipeline job that repeats the automated part.
The report needs a scope summary, the method and tool versions, the threat model, findings with severity and status, what was not tested and what you recommend. Attach transcripts to findings rather than pasting them into the body. State clearly that red teaming finds problems; it cannot prove their absence.
Automate the repeatable part. promptfoo’s CI documentation shows a GitHub Actions step with
promptfoo/promptfoo-action@v1and the type set to redteam. Run it on every change to prompts, tools or models and on a schedule, and run the manual sessions again when the system gains a new tool or data source. Article 55(1)(a) of the EU AI Act requires providers of general-purpose AI models with systemic risk to conduct and document adversarial testing; even if that does not apply to you, documenting it is good practice. Feed the results into how to audit AI systems evidence.GitHub Actions step from the promptfoo CI documentation · yaml - name: Run red team scan uses: promptfoo/promptfoo-action@v1 with: type: 'redteam' config: 'promptfooconfig.yaml' openai-api-key: ${{ secrets.OPENAI_API_KEY }} github-token: ${{ secrets.GITHUB_TOKEN }}Checked against: promptfoo: CI/CD integration, EU AI Act Article 55
What automated tools will not tell you
Automated suites find known patterns quickly. They do not understand your business. They will not notice that your refund tool accepts a negative amount, or that two users’ documents share an index. Those findings come from people who read the threat model and try to abuse the system the way a bad customer, a careless employee or a malicious document would. Budget at least as much human time as tool time, and do not report “no findings” from a tool run as “secure”.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| promptfoo redteam run reports nearly every attack as failed or errors on every call. | The target configuration does not match your endpoint: wrong URL, body shape, authentication header or transformResponse expression. | Send one request by hand with curl, copy the exact body and headers into the target, and set transformResponse to read the field your endpoint returns. |
| garak runs but nothing in the report fails, even on a model you expect to be weak. | The probe set is narrow, the detector did not match the output format, or the target name is wrong so the run used something unexpected. | Run with --list_probes to confirm what you selected, check the report file named at the start of the run, and widen the --spec selection. Remember a clean scan is not proof of safety. |
| pip install pyrit fails or installs an unexpected version. | Python version outside the supported range, or a conflict with other packages in the environment. | Use a fresh virtual environment with a supported Python (the docs recommend 3.13; PyPI lists 3.10 to below 3.15) and pin the version once it works. |
| Results differ every run and you cannot tell if a fix worked. | Attacks are generated and sampled randomly, and model outputs vary. | Keep a fixed set of confirmed attacks as regression cases, run them several times, and judge fixes by success rate, not a single pass. |
| Costs climb during a run. | Thousands of attack prompts, each also graded by a model. | Cap spend on the account, reduce numTests, drop strategies you do not need yet, use a smaller grader for first passes and a stronger one for final runs. |
| A tool reports a successful injection but a human cannot reproduce it. | The grader misjudged the output, or the attack depended on a state that no longer exists. | Replay the transcript against a fresh session; if it fails, mark it a false positive and add a note to the grader configuration or case. |
Verify it worked
Next steps
- How to stop prompt injection: The layered fixes for your top findings.
- How to validate your AI: Turn confirmed attacks into permanent regression gates.
- How to secure MCP servers: Test and harden the tools your agent can call.
- OWASP LLM Top 10 for engineers: A control for each of the ten risks.
- LLM red-teaming checklist: A shorter pre-deployment checklist.
Related guides
- How to Stop Prompt Injection: Layered Defences That Work: A layered defence for prompt injection: assume it will happen, keep untrusted content apart from instructions, give agents the fewest tools and shortest-lived privileges, require human approval for risky actions, block data leaving, and test with canary documents.
- How to Validate Your AI: Eval Sets, Gates, Evidence: A system-level method to validate an AI product: define the task and risk, build a held-out eval set, score it, gate releases, sample for human review, monitor and keep an evidence pack.
- How to Secure MCP Servers: OAuth, Scopes, Allow-Lists: Harden an MCP deployment against the attacks the specification names: validate token audience, never pass tokens through, ask for minimal scopes, sandbox local servers, allow-list servers, gate sensitive tools with a human, and log every call.
- How to Audit AI Systems: Scope, Evidence, Findings: How to audit an AI system, internally or for a client: scope it, choose criteria, request and sample evidence, test logs, change control and human oversight, and write findings that can be fixed.
- How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
Frequently asked questions
What is AI red teaming?
AI red teaming is authorised adversarial testing of an AI system: people and tools try to make it leak data, ignore its rules, take actions it should not or produce harmful output, so the weaknesses are found and fixed before real attackers find them. It covers the model, the application around it, its infrastructure and its behaviour at runtime.
What are the best open-source LLM red teaming tools?
Three are widely used and documented: promptfoo, which generates and runs attacks against an application and fits into CI; garak, NVIDIA’s scanner that probes a model across many attack families; and PyRIT, Microsoft’s framework for custom and multi-turn campaigns. They overlap but differ in focus, so most teams combine them with manual testing.
How often should we red team an LLM application?
Run the automated suite on every change to prompts, tools, models or retrieval data and on a schedule. Repeat manual sessions when the system gains a new tool, data source or user group, and after any incident. A single pre-launch exercise goes stale quickly because attacks and models both change.
Is red teaming required by the EU AI Act?
Article 55(1)(a) requires providers of general-purpose AI models with systemic risk to perform model evaluation including conducting and documenting adversarial testing. For high-risk AI systems, Article 15 requires resilience to attempts to alter use, outputs or performance by exploiting vulnerabilities. Check which duties apply to you and take legal advice.
Can I red team a model hosted by a third party?
Only within the provider’s terms and with the owner’s authorisation. Read the usage policy before running automated attack suites against a hosted API, and prefer testing your own application and its guardrails rather than the provider’s model. If the provider has a security testing process, use it.
How do I measure red team results?
Record the attack success rate over repeated attempts, the impact of each successful attack and the root cause. Score impact and reliability together, keep the bands stable between exercises and track the number of confirmed attacks turned into regression tests. Counts of prompts sent are not a measure of coverage.
How Swfte can help
Everything above works with open-source tools. If you want a security view of AI systems alongside governance, Swfte’s SecOps pages describe how it approaches red teaming and the OWASP list.
- AI red teaming: How Swfte frames adversarial testing.
- OWASP LLM Top 10: The list mapped to platform controls.
- Agent runtime security: Containing what an agent can do when fooled.
Swfte does not offer a managed red-team service on this page; if you need a human-led engagement, <red-team engagement availability - founder to fill>. You can complete every step above without Swfte.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- OWASP Top 10 for LLM Applications 2025: LLM01:2025 to LLM10:2025 titles
- OWASP LLM01:2025 Prompt Injection: Definition, direct and indirect injection, the seven prevention and mitigation strategies, note that RAG and fine-tuning do not fully mitigate
- OWASP GenAI Red Teaming Guide: Version 1.0, 22 January 2025; four areas: model evaluation, implementation testing, infrastructure assessment, runtime behavior analysis
- promptfoo: Red team quickstart: redteam setup / init --no-gui / run / report commands, Node.js 22.22.0 or newer, HTTP target example
- promptfoo: Red team configuration: targets and redteam top-level keys; purpose, plugins, strategies, numTests, provider fields
- promptfoo: OWASP LLM Top 10: owasp:llm and owasp:llm:01/02/06 presets; plugin mapping
- promptfoo: HTTP provider: url, method, headers, body with {{prompt}}, transformResponse
- promptfoo: CI/CD integration: promptfoo-action@v1 with type redteam
- NVIDIA garak README: pip install -U garak; Python 3.11 to 3.13 in the conda route; example commands; probe table; three log types
- garak CLI reference: --target_type, --target_name, --spec selectors, --list_probes, --generations, --parallel_attempts, --report_prefix
- PyRIT on PyPI: pip install pyrit; version 1.1.0 released 4 September 2026; Python 3.10 to below 3.15; MIT; Alpha status
- PyRIT documentation (microsoft.github.io/PyRIT): Python 3.13 recommended; ~/.pyrit/.env and ~/.pyrit/.pyrit_conf; Scanner, GUI and Framework modes
- Simon Willison: The lethal trifecta for AI agents: Three conditions, 16 June 2025
- Meta: Agents Rule of Two: Properties A, B, C and the rule, 31 October 2025
- EU AI Act Article 55: Article 55(1)(a): adversarial testing for GPAI models with systemic risk
Topics
- red teaming
- security
- OWASP LLM Top 10
- promptfoo
- garak
- PyRIT
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-red-team-an-llm.