SecOps / Security for AI

Prompt injection defense for agentic systems

Assume the model will sometimes be fooled. Build the system so that being fooled does not become a breach.

Prompt injection is the top-ranked risk in the OWASP Top 10 for LLM Applications because the model cannot reliably tell instructions from data. In a chatbot that is an embarrassment. In an agent that can read mail, call tools and change records, it is an access-control problem, and it needs an access-control answer.

The problem: instructions and data share one channel

A language model reads everything as text. The system prompt, the user's request, the web page it fetched, the email it summarised and the tool output it received all arrive in the same stream. An attacker who can put words in front of the model can try to give it orders. This is direct prompt injection when the attacker is the user, and indirect prompt injection when the instructions are hidden in content the agent retrieves: a document, a calendar invite, a ticket, a web page, a tool description.

Indirect injection is the version that matters for agents. The user did nothing wrong; the agent simply read something. Public research on zero-click exploits of production assistants showed the pattern: a crafted email or document, an agent with access to private data, and an outbound channel for the data to leave. None of those three ingredients is a bug on its own. Together they are a breach.

Filters and detectors help, and we use them, but they are probabilistic and adaptive attackers get past them. The durable defense is architectural. Meta's published "Agents Rule of Two" states the idea plainly: within one session an agent should combine no more than two of three properties, which are processing untrustworthy input, accessing sensitive systems or data, and changing state or communicating externally. If it needs all three, it should not run unsupervised.

How Swfte limits what an injected agent can do

The platform treats prompt injection as a layered control problem. Detection is one layer; permissions, approvals, egress and evidence are the others.

  • Least-privilege tools

    Each agent has an explicit tool allowlist and scoped credentials. An injected instruction to "email this file to an outside address" fails when the agent has no email-send tool, or when sending is a gated action.

  • Approval gates on state change

    Actions that change state or reach outside the boundary, such as sending, paying, deleting and publishing, can require a named human to approve. This is the Rule of Two applied as policy.

  • Untrusted content is labelled as data

    Retrieved documents, tool output and external messages are passed to the model as untrusted data, with structure that keeps them apart from instructions. This reduces the success rate of injection; it does not remove it, which is why the other layers exist.

  • Egress and destination control

    Outbound calls from an agent run through the gateway, where destinations can be allowed or denied and data classes filtered. Exfiltration through a rendered link or a tool call meets a policy check first.

  • In-flight policy at the tool call

    Policy is evaluated before a tool call executes, not after. Nexus does this today for coding agents: protected files, denied commands and unapproved installs are blocked with the reason recorded.

  • Detection and evidence

    Suspected injection attempts are flagged and the full chain is kept: the content, the model output, the tool call and the policy decision. That is what lets a SOC investigate afterwards.

Where it sits on the platform

One platform with several entry points. Each product below plays a defined part in this capability.

Platform products and the role each plays
Platform entry pointRole in this capability
NexusEvaluates policy before each tool call and records blocked actions with the reason. Strongest today for coding agents.
StudioDefine each agent's tools, scopes and approval gates at build time, so the safe configuration is the default.
BuildXThe model gateway: one place to apply input and output screening and to log every model call.
ConnectScoped connections to external systems, so an agent holds the narrowest credential that does the job.
CortexControlled retrieval. Knowledge reaches agents through access rules rather than an open index.

What an agent that reads untrusted content can and cannot do

A Trust Profile for an inbox-and-documents assistant, the classic injection target. The limits are what contain a successful injection.

Inbox Triage Agent

Reads incoming mail and attachments, summarises and drafts replies.

Inbox Triage Agent: what it can do, cannot do, requires approval for, and records
CanCannotRequires approvalRecords
  • Read the mailbox and linked approved documents
  • Summarise and classify messages
  • Draft replies for a person to review
  • Look up approved knowledge-base articles
  • Send mail on its own to external addresses
  • Forward attachments or files outside the organisation
  • Follow links or call tools not on its allowlist
  • Change its own instructions or permissions
  • Sending any message externally
  • Opening or fetching a new external destination
  • Any access beyond the mailbox it is assigned
  • Message and attachment read
  • Model used and output
  • Tools called and any blocked call
  • Policy applied and reason
  • Approver and outcome

Controlled autonomy against injection risk

Exposure to untrusted content is a reason to hold autonomy lower, not higher. The level follows the worst thing an injected agent could do.

  1. L1 Assist

    Default for any agent that reads external content and holds sensitive access. It recommends; a person acts.

  2. L2 Approve

    The agent prepares outbound or state-changing actions; a person approves each one. The Rule of Two in practice.

  3. L3 Supervise

    Acts within narrow limits where the action cannot move data outside the boundary, with monitoring and alerts.

  4. L4 Autonomous

    Reserved for agents with no untrusted input, or no sensitive access, or no external effect, and a record to show it.

  5. L5 Adaptive

    Not applicable to agents exposed to untrusted content without gated, tested changes.

The five-level model is described in full on the controlled autonomy page.

Frameworks the design is mapped to

These mappings show which published risks each control is intended to address. They are not an assessment result.

Framework entries and how the controls are designed to address them
FrameworkEntryHow the design addresses it
OWASP Top 10 for LLM Applications (2025)LLM01 Prompt InjectionConstrained tool scope, untrusted-content handling, output checks and human approval on high-risk actions.
OWASP Top 10 for LLM Applications (2025)LLM05 Improper Output HandlingModel output is validated before it reaches a tool, a shell or a downstream system.
OWASP Top 10 for Agentic Applications (2026)ASI01 Agent Goal HijackGoals and permissions are set outside the model's context; retrieved content cannot rewrite them.
MITRE ATLASAML.T0051 LLM Prompt InjectionDetection signals and the evidence chain are shaped so a SOC can follow an injection attempt through ATLAS techniques.

EU angle: robustness and security you can show

These are the instruments we most often see cited. Swfte supports evidence; it does not certify your use.

EU instruments and what Swfte supports evidence for
InstrumentReferenceSupports evidence for
EU AI ActArt. 15 accuracy, robustness and cybersecurityArt. 15 expects high-risk systems to resist attempts by third parties to alter their use or outputs. Supports evidence of the controls and tests you run against manipulation.
GDPRArt. 32 security of processingSupports evidence of technical measures that stop an injected agent from moving personal data out of its boundary.
NIS2Art. 21 risk-management measuresSupports evidence of access control, supply-chain and incident-handling measures for in-scope entities.

Swfte is built compliance-by-design. It provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. Nothing on this page is legal advice. Swfte's own security attestations are listed on the trust page.

What this page does not claim

  • No product can promise to block all prompt injection. Anyone claiming a detection rate is selling a filter, and filters are bypassed. We publish no such rate.
  • Nexus enforcement is deepest today for coding agents. Coverage for other agent types is designed in and varies by connector.
  • Model-level defenses improve over time and are not a substitute for permissions and approvals.
  • Nothing here claims that your use of the platform is compliant with any regulation. Swfte's own security attestations are listed on the trust page.

Frequently asked questions

What is the difference between direct and indirect prompt injection?

In direct injection the person typing to the model is the attacker. In indirect injection the attacker plants instructions in content the agent later retrieves, such as a web page, an email, a document or a tool description. Indirect injection is the larger risk for agents because the legitimate user never sees the attack.

Can a prompt injection filter stop it completely?

No. Filters and classifiers catch known patterns and lower the rate, but attackers adapt. The reliable defenses are limiting what the agent can do, requiring approval for risky actions and controlling where data can go, so that a successful injection has little to reach.

What is the Agents Rule of Two?

A guideline published by Meta: in a single session an agent should have at most two of three properties. It processes untrustworthy input, it accesses sensitive systems or private data, and it can change state or communicate externally. If all three are required, a human should approve, or the work should be split across sessions.

Does Swfte block prompt injection?

Swfte provides layered controls: scoped tools, approval gates, egress policy, screening at the model gateway and an evidence trail. Policy is evaluated before a tool call executes. We do not claim it blocks every attempt, and we do not publish a detection rate.

How do we test our own agents?

With adversarial testing mapped to OWASP and MITRE ATLAS, run in a sandbox with the same policy as production. See the AI red teaming page.

The rest of SecOps

Nine pages cover both halves. Each has a different job.

Build Prompt injection defense with Swfte

Start with one agent and one policy, or talk to the team about your environment, your data and your regulators.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.