Agentic retrieval
RAG agent: how it differs from RAG and the controls it needs
The extra failure modes of a RAG agent and the controls that bound them: step limits, budgets, approvals and trace logs, with what Swfte has built for each.
A RAG agent is a language model that decides for itself whether to retrieve, what to search for, whether the results are good enough and whether to try again. That gives it reach a fixed pipeline lacks. It also adds loops, runaway cost, tool misuse and more routes for injected text. The controls are step limits, spend budgets, approval gates and a full trace of each run.
Last verified 2026-10-07. Sources are listed at the end of the page.
What is a RAG agent?
Anthropic’s “Building effective agents” (19 December 2024) separates two designs. Workflows are “systems where LLMs and tools are orchestrated through predefined code paths.” Agents are “systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.” A standard RAG pipeline is a workflow. A RAG agent is the same retrieval placed under an agent’s control.
The LangGraph agentic RAG tutorial shows the shape. The model decides whether to call a retriever tool or answer directly. After retrieval, a grading step scores whether the documents are relevant. If they are not, the question is rewritten and retrieval is tried again. If they are, the model answers from them. The loop is the point, and the loop is also the risk.
How is this page different from the other agentic RAG pages?
| Page | What it covers |
|---|---|
| Agentic RAG | The pattern, how it differs from a fixed pipeline, its architecture and cost per query. |
| RAG vs MCP | When to retrieve from an index and when to give an agent live tool access. |
| RAG pipeline | The stages of the fixed pipeline that an agent’s retrieval tool still depends on. |
| This page | What goes wrong when an agent controls retrieval, and the controls that bound it, with Swfte’s status for each. |
What extra failure modes does an agent add to RAG?
| Failure mode | What happens | Control |
|---|---|---|
| Loop | The grader keeps rejecting results, the question is rewritten again and again, and nothing converges. | A maximum number of iterations. Anthropic recommends stopping conditions such as this. |
| Runaway cost | Every rewrite, retrieval and check is another model call. Anthropic reports that agents typically use about 4 times more tokens than chat and multi-agent systems about 15 times more. | A token or spend budget per run and per workspace, with an alert before the cap. |
| Tool misuse | The agent calls a tool with the wrong arguments, or one with side effects it should not have had. | Least privilege for each tool, typed inputs and approval before any write. |
| Injection | Retrieved text carries instructions. LangChain’s documentation says no prompt or delimiter strategy fully prevents this. | Treat retrieved text as data, restrict the agent’s tools and validate outputs. OWASP lists human approval for sensitive operations. |
| Compounding errors | A weak first retrieval shapes every later step. Anthropic notes that agents bring a potential for compounding errors. | Grade retrieval at each step, log it, and stop for a person when confidence stays low. |
| Permission widening | The agent searches with a broad service identity and returns what the asking person could not open. | Retrieve as the person the agent acts for, filtered inside the query. |
Which controls should a RAG agent have before it goes live?
Step limit
Set a hard maximum of retrieval and rewrite steps. When it is reached the agent answers with what it has, says retrieval was incomplete, or hands over to a person.
Budget
Set a token or spend ceiling for each run and for the workspace. Alert before the ceiling. A limit that only logs after the money is gone is not a control.
Approval gates
Require a person to approve any action that changes something outside the agent: sending, writing, paying, deleting. LangGraph’s interrupt pattern pauses a run at such a point and resumes on the person’s decision. See human in the loop.
Trace logging
Record each step: the query as written and as rewritten, the filter, the chunks returned, the grade, the tool calls and the final answer. Without the trace you cannot tell a retrieval problem from a model problem.
Sandboxed tests
Anthropic recommends extensive testing in sandboxed environments along with guardrails. Test with hostile documents in the corpus before real users arrive.
When should you not build a RAG agent?
Anthropic says agents suit open-ended problems where the number of steps cannot be predicted, and that their autonomy means higher costs. If your questions are lookups against one corpus, a fixed pipeline is cheaper, easier to test and has no loop to bound. Add the agent when a measured set of failures needs it: questions that span sources, or that need a second search after the first.
Where Swfte fits: what is built for a RAG agent
The existing page agents on the brain holds the full status table. In short:
| Part | Status | Note |
|---|---|---|
| Graph of people, groups and reporting lines the agent can read | Built | Through the appliance’s local API, scoped by token. |
| Permission-filtered search over documents | In progress | Libraries exist. Endpoints are not finished. |
| Context-package API and MCP server for agents | Roadmap | Designed for. Not built. |
| Agents built in Studio; Nexus capturing agent actions, enforcing policy and tracing them | Built | Not yet wired to the company brain. |
| Spend and token caps in Connect | Built | Workspace and per-model caps, with alerts and an action when a cap is hit, such as downgrading to a cheaper model. They bound cost. They do not count steps. |
| Human approval before an action | Built | Several approval mechanisms exist in separate parts of the platform. See human in the loop. |
A step limit belongs in the agent definition or the framework you run. This page does not state that Studio exposes one. If your agent is a single fixed pipeline over one corpus, you do not need Swfte for this: a framework such as LangGraph and your own limits will do. Governance policy is described on platform governance.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Anthropic, Building effective agents. Workflow and agent definitions, the cost trade-off, stopping conditions, sandboxed testing and human checkpoints.
- LangGraph, agentic RAG tutorial. The retrieve, grade, rewrite and generate loop with conditional edges.
- LangChain, build a RAG agent. The statement that no prompt or delimiter strategy fully prevents indirect prompt injection.
- Anthropic, How we built our multi-agent research system. Token multipliers for agents and multi-agent systems, and compounding errors.
- OWASP, LLM01 Prompt Injection. Indirect prompt injection and the mitigations, including human approval for sensitive operations.
- LangGraph, interrupts. Pausing a run for human approval before a critical action and resuming with the decision.
Frequently asked questions
What is a RAG agent?
A RAG agent is a language model that controls its own retrieval. It decides whether to search, writes and rewrites the query, grades what comes back and loops until the evidence is good enough or a limit is reached. A standard RAG pipeline runs the same fixed steps for every question.
What is the difference between RAG and a RAG agent?
In RAG, code fixes the steps: retrieve, then generate. In a RAG agent the model chooses the steps, which Anthropic describes as dynamically directing its own process and tool use. That helps with multi-step questions and costs more, with a risk of loops and compounding errors.
How do you stop a RAG agent from looping?
Set a maximum number of iterations and decide in advance what happens at the limit: answer with partial evidence, say retrieval was incomplete, or escalate to a person. Add a spend budget as a second stop. Log every rewrite so you can see why the loop did not converge.
Does a RAG agent make prompt injection worse?
It gives injection more to work with. An agent reads more retrieved text and holds tools that can act on it. LangChain’s documentation says no prompt or delimiter strategy fully prevents indirect injection, so limit the agent’s tools, retrieve as the asking person and require approval before any write.
Is a RAG agent more expensive than RAG?
Usually, per question, because each rewrite, retrieval and check is another model call. Anthropic reports that agents typically use about 4 times more tokens than chat and multi-agent systems about 15 times more. That figure is for its own research system. Measure your own agent’s cost per answered question.
Can Swfte agents read the company brain today?
Not through a finished agent interface. The graph and its local API are built. Permission-filtered document search is in progress. The context-package API and MCP server are on the roadmap, as is wiring Studio and Nexus into the brain. See the agents on the brain page for the full table.