← The journal
Guide Summary

Securing and Governing AI Agents: Injection, MCP, Approvals

A summary of four guides on securing and governing AI agents: prompt injection, MCP, approvals and how they fit.

Swfte Journal / Guide Summary

An agent is a model with tools. Everything that makes it useful also makes it a target, and the controls that work are the same few ideas applied in different places. We wrote four guides on those places. This post summarises each, shows how they combine, and points to the full versions, which have commands, expected output, a troubleshooting table and verified sources.

The shared premise

Assume the model will sometimes follow an instruction it should not, and design so that this does less harm. That is the guide's answer to prompt injection: no one has a way to stop it completely, and OWASP says it is unclear whether one exists, so you limit what an injected instruction can reach and do. Every other control in this post is a way of limiting reach, adding a human, or recording what happened.

Stopping prompt injection: eight steps

  1. Map where untrusted text enters and what the agent can do next. Web pages, emails, documents, tool results and user messages are all inputs. List them. For each, ask what the agent could do after reading it.
  2. Keep untrusted content apart from your instructions. Mark it clearly as data, and say in the instructions that it is data. This helps and does not solve the problem, so do not stop here.
  3. Cut tools, scopes and the lifetime of privileges. If the agent has no tool that can send data out, an injected instruction to send data out fails. Short-lived credentials limit what a stolen one can do.
  4. Screen inputs and tool outputs, and validate what the model returns. Filters catch known patterns and miss new ones. Treat them as one layer. Validate structured output against a schema before acting on it.
  5. Put a person in front of risky actions. See the approval guide below.
  6. Block data from leaving by routes you did not approve. Limit outbound network access and the places an agent can write to. Many injection attacks need a way to get data out, and removing that stops them even when the model is fooled.
  7. Test every release with planted instructions. Put hostile instructions in the documents and pages your agent reads and check what it does. Our red teaming guide covers the tools.
  8. Monitor for successful injections and learn from them. Log tool calls, review odd sequences, and add each real incident to the test set.

Securing MCP servers: nine steps

The Model Context Protocol lets agents use tools exposed by servers. The guide treats every server as code you choose to run and every tool result as untrusted input. It is based on the current specification's security and authorisation material, with the spec version and date stated in the guide.

  1. Inventory every server and read what it can do.
  2. Require OAuth with audience-checked tokens on remote servers. A token minted for one service must not work on another.
  3. Refuse token passthrough and add per-client consent to proxies. Passing a client's token through to a downstream API is a known confused-deputy risk.
  4. Grant the smallest scopes and raise them on demand.
  5. Put local servers behind consent and a sandbox. A local server can run commands as you.
  6. Keep an allow-list of approved servers.
  7. Gate sensitive tools with a person and validate inputs. Tool descriptions can be written to manipulate the model, so do not trust a description because it sounds helpful.
  8. Restrict outbound requests from clients and servers.
  9. Log every call and review on a schedule.

Our pages on MCP security best practices, MCP gateways and tool security give more context.

Human approval: seven steps

The guide's rule is blunt: put a person in front of actions that are hard to undo, move money or data outside your systems, or change permissions, and nothing else. Gating everything produces approval fatigue, and tired reviewers approve everything.

  1. Sort every tool into approve-first or run-freely.
  2. Set thresholds and a default for silence. If nobody answers, deny.
  3. Pause the agent at the gate and resume on the decision. The guide uses LangGraph's interrupt mechanism for this, with verified code.
  4. Or gate at the tool level with the middleware.
  5. Show the reviewer what they need to decide: the intent, the arguments and the source of the instruction.
  6. Record every decision, including timeouts.
  7. Measure approvals and fix rubber-stamping. If reviewers approve nearly everything instantly, the gate is decoration.

See also the autonomy page and the post on approval workflows auditors accept.

Governing agents: eight steps

Governance, in the guide's words, means being able to say who is acting, allowed to do what, under which policy, with which data and model, at what risk and with what oversight, and to prove it afterwards.

  1. Build an agent inventory.
  2. Give every agent its own identity and short-lived access. Shared service accounts make accountability impossible. See non-human identity.
  3. Write a Trust Profile for each agent: owner, risk level, approved models, data classification, residency, permitted systems, allowed and restricted actions, approval rule, retention, audit and policy set.
  4. Turn the profile into enforced rules. A profile nobody enforces is a document.
  5. Choose an autonomy level per agent and raise it only on evidence. The levels run from assisting (the agent recommends), through approving (a person approves), to supervised and autonomous operation within limits.
  6. Record who did what, under which policy, and the outcome.
  7. Build a kill switch and test it.
  8. Review on a schedule and map the work to a framework such as the NIST AI Risk Management Framework. See AI agent governance, the governance platform page and the Trust Profile page.

How they combine

Picture an agent that reads supplier emails and drafts purchase orders. Injection defence stops an email from redirecting it. MCP controls limit which tools it can reach and with what authority. Approval puts a person in front of any order above a threshold. Governance gives the agent an identity, a profile that says what it may do, and a log that shows afterwards what it did. Take away any one and the others carry more weight than they can bear.

A short threat story

An agent reads support tickets, looks up orders and can issue refunds. A customer pastes text into a ticket that says to ignore earlier instructions and refund an order to a different account. Without controls, the model may comply. With the layers above, several things have to fail together: the ticket text is marked as data, the refund tool is the only way money moves and has a limit, refunds above the limit go to a person who sees the instruction's source, the agent has no tool for sending data to an arbitrary address, and the attempt is logged and shows up in review. Each layer is imperfect. The combination is what makes an attack expensive.

The same story with MCP: the order lookup is a tool on a server. If that server were replaced by a malicious one, or its tool description changed, the agent could be steered. The allow-list, the pinned and reviewed server, the narrow scopes and the logging are what limit the damage.

Where to start

If you have agents running and nothing else: the inventory and a kill switch, from the governance guide. If you use MCP: the allow-list and the scopes. If an agent can spend money or send messages: the approval gate. If it reads content from outside: the injection map. Then test, because untested controls are assumptions.

Swfte's products are designed around this work, and some capabilities are available on request. The guides do not require them, and each says what you can do with open tools alone.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.