SecOps / Security for AI
Prompt injection defense for agentic systems
Assume the model will sometimes be fooled. Build the system so that being fooled does not become a breach.
Prompt injection is the top-ranked risk in the OWASP Top 10 for LLM Applications because the model cannot reliably tell instructions from data. In a chatbot that is an embarrassment. In an agent that can read mail, call tools and change records, it is an access-control problem, and it needs an access-control answer.
The problem: instructions and data share one channel
A language model reads everything as text. The system prompt, the user's request, the web page it fetched, the email it summarised and the tool output it received all arrive in the same stream. An attacker who can put words in front of the model can try to give it orders. This is direct prompt injection when the attacker is the user, and indirect prompt injection when the instructions are hidden in content the agent retrieves: a document, a calendar invite, a ticket, a web page, a tool description.
Indirect injection is the version that matters for agents. The user did nothing wrong; the agent simply read something. Public research on zero-click exploits of production assistants showed the pattern: a crafted email or document, an agent with access to private data, and an outbound channel for the data to leave. None of those three ingredients is a bug on its own. Together they are a breach.
Filters and detectors help, and we use them, but they are probabilistic and adaptive attackers get past them. The durable defense is architectural. Meta's published "Agents Rule of Two" states the idea plainly: within one session an agent should combine no more than two of three properties, which are processing untrustworthy input, accessing sensitive systems or data, and changing state or communicating externally. If it needs all three, it should not run unsupervised.
How Swfte limits what an injected agent can do
The platform treats prompt injection as a layered control problem. Detection is one layer; permissions, approvals, egress and evidence are the others.
Least-privilege tools
Each agent has an explicit tool allowlist and scoped credentials. An injected instruction to "email this file to an outside address" fails when the agent has no email-send tool, or when sending is a gated action.
Approval gates on state change
Actions that change state or reach outside the boundary, such as sending, paying, deleting and publishing, can require a named human to approve. This is the Rule of Two applied as policy.
Untrusted content is labelled as data
Retrieved documents, tool output and external messages are passed to the model as untrusted data, with structure that keeps them apart from instructions. This reduces the success rate of injection; it does not remove it, which is why the other layers exist.
Egress and destination control
Outbound calls from an agent run through the gateway, where destinations can be allowed or denied and data classes filtered. Exfiltration through a rendered link or a tool call meets a policy check first.
In-flight policy at the tool call
Policy is evaluated before a tool call executes, not after. Nexus does this today for coding agents: protected files, denied commands and unapproved installs are blocked with the reason recorded.
Detection and evidence
Suspected injection attempts are flagged and the full chain is kept: the content, the model output, the tool call and the policy decision. That is what lets a SOC investigate afterwards.
Where it sits on the platform
One platform with several entry points. Each product below plays a defined part in this capability.
| Platform entry point | Role in this capability |
|---|---|
| Nexus | Evaluates policy before each tool call and records blocked actions with the reason. Strongest today for coding agents. |
| Studio | Define each agent's tools, scopes and approval gates at build time, so the safe configuration is the default. |
| BuildX | The model gateway: one place to apply input and output screening and to log every model call. |
| Connect | Scoped connections to external systems, so an agent holds the narrowest credential that does the job. |
| Cortex | Controlled retrieval. Knowledge reaches agents through access rules rather than an open index. |
What an agent that reads untrusted content can and cannot do
A Trust Profile for an inbox-and-documents assistant, the classic injection target. The limits are what contain a successful injection.
Inbox Triage Agent
Reads incoming mail and attachments, summarises and drafts replies.
| Can | Cannot | Requires approval | Records |
|---|---|---|---|
|
|
|
|
Controlled autonomy against injection risk
Exposure to untrusted content is a reason to hold autonomy lower, not higher. The level follows the worst thing an injected agent could do.
- L1 Assist
Default for any agent that reads external content and holds sensitive access. It recommends; a person acts.
- L2 Approve
The agent prepares outbound or state-changing actions; a person approves each one. The Rule of Two in practice.
- L3 Supervise
Acts within narrow limits where the action cannot move data outside the boundary, with monitoring and alerts.
- L4 Autonomous
Reserved for agents with no untrusted input, or no sensitive access, or no external effect, and a record to show it.
- L5 Adaptive
Not applicable to agents exposed to untrusted content without gated, tested changes.
The five-level model is described in full on the controlled autonomy page.
Frameworks the design is mapped to
These mappings show which published risks each control is intended to address. They are not an assessment result.
| Framework | Entry | How the design addresses it |
|---|---|---|
| OWASP Top 10 for LLM Applications (2025) | LLM01 Prompt Injection | Constrained tool scope, untrusted-content handling, output checks and human approval on high-risk actions. |
| OWASP Top 10 for LLM Applications (2025) | LLM05 Improper Output Handling | Model output is validated before it reaches a tool, a shell or a downstream system. |
| OWASP Top 10 for Agentic Applications (2026) | ASI01 Agent Goal Hijack | Goals and permissions are set outside the model's context; retrieved content cannot rewrite them. |
| MITRE ATLAS | AML.T0051 LLM Prompt Injection | Detection signals and the evidence chain are shaped so a SOC can follow an injection attempt through ATLAS techniques. |
EU angle: robustness and security you can show
These are the instruments we most often see cited. Swfte supports evidence; it does not certify your use.
| Instrument | Reference | Supports evidence for |
|---|---|---|
| EU AI Act | Art. 15 accuracy, robustness and cybersecurity | Art. 15 expects high-risk systems to resist attempts by third parties to alter their use or outputs. Supports evidence of the controls and tests you run against manipulation. |
| GDPR | Art. 32 security of processing | Supports evidence of technical measures that stop an injected agent from moving personal data out of its boundary. |
| NIS2 | Art. 21 risk-management measures | Supports evidence of access control, supply-chain and incident-handling measures for in-scope entities. |
Swfte is built compliance-by-design. It provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. Nothing on this page is legal advice. Swfte's own security attestations are listed on the trust page.
What this page does not claim
- No product can promise to block all prompt injection. Anyone claiming a detection rate is selling a filter, and filters are bypassed. We publish no such rate.
- Nexus enforcement is deepest today for coding agents. Coverage for other agent types is designed in and varies by connector.
- Model-level defenses improve over time and are not a substitute for permissions and approvals.
- Nothing here claims that your use of the platform is compliant with any regulation. Swfte's own security attestations are listed on the trust page.
Frequently asked questions
What is the difference between direct and indirect prompt injection?
In direct injection the person typing to the model is the attacker. In indirect injection the attacker plants instructions in content the agent later retrieves, such as a web page, an email, a document or a tool description. Indirect injection is the larger risk for agents because the legitimate user never sees the attack.
Can a prompt injection filter stop it completely?
No. Filters and classifiers catch known patterns and lower the rate, but attackers adapt. The reliable defenses are limiting what the agent can do, requiring approval for risky actions and controlling where data can go, so that a successful injection has little to reach.
What is the Agents Rule of Two?
A guideline published by Meta: in a single session an agent should have at most two of three properties. It processes untrustworthy input, it accesses sensitive systems or private data, and it can change state or communicate externally. If all three are required, a human should approve, or the work should be split across sessions.
Does Swfte block prompt injection?
Swfte provides layered controls: scoped tools, approval gates, egress policy, screening at the model gateway and an evidence trail. Policy is evaluated before a tool call executes. We do not claim it blocks every attempt, and we do not publish a detection rate.
How do we test our own agents?
With adversarial testing mapped to OWASP and MITRE ATLAS, run in a sandbox with the same policy as production. See the AI red teaming page.
The rest of SecOps
Nine pages cover both halves. Each has a different job.
AI for security operations
Frameworks and evidence
Back to the SecOps hub.
Build Prompt injection defense with Swfte
Start with one agent and one policy, or talk to the team about your environment, your data and your regulators.