SecOps / Security for AI

AI red teaming: test your agents before an attacker does

Run adversarial tests against your own models, agents and tool chains, map each finding to a published framework and retest after every fix.

A control you have not attacked is a control you have only hoped for. AI red teaming applies adversarial testing to the AI layer: prompt injection, data exfiltration through tools, goal hijack, memory poisoning and excessive agency. The output is not a score; it is a list of reproducible findings with an owner and a retest.

The problem: AI behaviour changes, and so does the attack surface

Traditional penetration testing assumes a system that behaves the same way each time. A language-model application does not. The same input can yield different outputs, a model update can shift behaviour overnight, and the attack surface includes natural language, retrieved documents, tool descriptions and memory. A test that passed last quarter says little about this one.

Many teams therefore test once before launch and never again, or rely on a vendor's published safety evaluations, which test the model and not your application. Your risk lives in the combination: your prompts, your data, your tools, your permissions. Only a test against that combination tells you something useful.

Red teaming agents adds a further dimension. The question is not only whether the model says something it should not, but whether a manipulated agent can take an action it should not: read a file, call a tool, send a message. That is a question about the permissions and policy around the model, and it can be answered by trying.

How testing runs on the platform

Red teaming is itself a governed workflow. Attackers are agents, the target runs in a sandbox, and a person approves the scope.

  • Scope and authorisation first

    A named owner approves what is in scope: which agents, tools, data and environments. Testing does not start without it, and out-of-scope targets are blocked by policy, not by good intentions.

  • Sandboxed targets

    Tests run against staging or sandboxed copies with synthetic data and the same policy as production, so a finding reflects how the real control behaves and nothing real is at risk.

  • Framework-mapped test sets

    Test cases are organised against the OWASP Top 10 for LLM Applications, the OWASP Top 10 for Agentic Applications and MITRE ATLAS techniques, so coverage is visible and gaps are named.

  • Reproducible findings

    Each finding records the input, the context, the model and version, the tool calls it triggered and the policy decision. Someone else can reproduce it, which is what a fix needs.

  • Fix, then retest

    A finding closes only when the retest passes under the same conditions. Because behaviour drifts, tests are rerun on model, prompt or tool changes.

  • Findings feed the controls

    A successful attack becomes a policy rule, an approval gate or a detection, so the system gets harder rather than the report getting longer.

Where it sits on the platform

One platform with several entry points. Each product below plays a defined part in this capability.

Platform products and the role each plays
Platform entry pointRole in this capability
StudioBuild test-harness agents and workflows, with approval gates on scope and on any step that leaves the sandbox.
NexusCaptures the tool actions and blocked calls during a test, which is the evidence for each finding.
BuildXTest across approved models through one gateway, and compare behaviour when a model changes.
CortexSeed the sandbox with synthetic, access-controlled knowledge to test retrieval and leakage paths.
Trust FabricEvidence and traceability for test runs, so results stand up in a review.

What a red-team agent can and cannot do

A Trust Profile for an attacker agent. The limits matter more than the capabilities: an offensive agent is dangerous precisely where its scope is loose.

Red-Team Agent

Generates and runs adversarial test cases against an approved target in a sandbox.

Red-Team Agent: what it can do, cannot do, requires approval for, and records
CanCannotRequires approvalRecords
  • Generate injection, exfiltration and goal-hijack test cases
  • Run them against in-scope sandboxed targets
  • Record inputs, outputs, tool calls and policy decisions
  • Draft a finding with a suggested control
  • Test any target outside the approved scope
  • Use real customer data or production credentials
  • Reach the open internet from the sandbox
  • Close or downgrade a finding itself
  • Expanding scope or adding a target
  • Any test with a destructive or high-volume pattern
  • Publishing or sharing findings outside the owning team
  • Scope approval and approver
  • Test case, target, model and version
  • Tool calls and policy decisions
  • Result and reproduction steps
  • Retest outcome

Controlled autonomy for offensive testing

Offensive capability stays at low autonomy by default. Scope and impact set the level, and human approval on scope never goes away.

  1. L1 Assist

    The agent suggests test cases from the framework lists. A tester runs them.

  2. L2 Approve

    The agent prepares and queues runs; a person approves each batch against the scope.

  3. L3 Supervise

    The agent runs approved test sets within the sandbox and rate limits, monitored, and drafts findings.

  4. L4 Autonomous

    Regression suites rerun on every model, prompt or tool change within strict policy, and report results.

  5. L5 Adaptive

    Attack generation improves from past findings, with gated, versioned changes. Scope cannot be self-modified.

The five-level model is described in full on the controlled autonomy page.

Frameworks the test sets are organised against

Organising tests against published lists makes coverage visible. It does not mean the lists are exhaustive.

Framework entries and how the controls are designed to address them
FrameworkEntryHow the design addresses it
OWASP Top 10 for LLM Applications (2025)LLM01 to LLM10Test cases tagged to each risk, from prompt injection to unbounded consumption, so a gap in coverage is a named gap.
OWASP Top 10 for Agentic Applications (2026)ASI01 to ASI10Goal hijack, tool misuse, memory poisoning, inter-agent communication and rogue-agent scenarios for multi-step agents.
MITRE ATLASTactics and techniques, for example AML.T0051 LLM Prompt InjectionFindings are labelled with ATLAS techniques so the SOC and the red team share one vocabulary. Confirm current IDs at atlas.mitre.org; the matrix is versioned and changes often.

EU angle: testing as evidence of robustness

Several instruments expect that you test the resilience of what you run. Test records are the evidence.

EU instruments and what Swfte supports evidence for
InstrumentReferenceSupports evidence for
EU AI ActArt. 15 accuracy, robustness and cybersecurity; Art. 9 risk managementSupports evidence of testing against manipulation and of how findings fed risk treatment, for systems where those articles apply.
DORADigital operational resilience testingSupports test records for financial entities in scope. Whether any activity counts as a required test is for your supervisor and testing provider.
NIS2Art. 21 policies to assess the effectiveness of measuresSupports evidence that you assess whether your AI security measures work.
GDPRArt. 32(1)(d) regular testing of security measuresSupports a record of regular testing of measures protecting personal data in AI systems.

Swfte is built compliance-by-design. It provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. Nothing on this page is legal advice. Swfte's own security attestations are listed on the trust page.

What this page does not claim

  • AI red teaming finds problems; it cannot prove their absence. A clean run is a data point, not an assurance.
  • Automated testing complements, and does not replace, human specialists and independent assessors.
  • We state no pass rate, vulnerability count or customer finding, because none is published here.
  • Nothing here claims that your use of the platform is compliant with any regulation. Swfte's own security attestations are listed on the trust page.

Frequently asked questions

What is AI red teaming?

It is adversarial testing of AI systems: attempting prompt injection, data exfiltration, goal hijack, tool misuse and similar attacks against your own models and agents, to find weaknesses before an attacker does. The result is a set of reproducible findings, each with an owner and a retest.

How is it different from a model safety evaluation?

A safety evaluation tests a model in isolation. Red teaming tests your application: your prompts, data, tools and permissions together. Risk lives in that combination.

Is it safe to let an agent attack our systems?

Only with controls: a named approver for scope, sandboxed targets with synthetic data, network limits and a record of every run. Offensive agents are governed agents, and their Trust Profile should state what they cannot do as clearly as what they can.

How often should we test?

On every change that can shift behaviour: a new model or model version, a prompt change, a new tool or permission, and a new data source. Plus a periodic broader exercise. Because behaviour drifts, a one-time test ages quickly.

Which frameworks do you use?

Test cases are organised against the OWASP Top 10 for LLM Applications (2025), the OWASP Top 10 for Agentic Applications (2026) and MITRE ATLAS techniques.

The rest of SecOps

Nine pages cover both halves. Each has a different job.

Build AI red teaming with Swfte

Start with one agent and one policy, or talk to the team about your environment, your data and your regulators.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.