Agentic retrieval

RAG agent: how it differs from RAG and the controls it needs

The extra failure modes of a RAG agent and the controls that bound them: step limits, budgets, approvals and trace logs, with what Swfte has built for each.

A RAG agent is a language model that decides for itself whether to retrieve, what to search for, whether the results are good enough and whether to try again. That gives it reach a fixed pipeline lacks. It also adds loops, runaway cost, tool misuse and more routes for injected text. The controls are step limits, spend budgets, approval gates and a full trace of each run.

Last verified 2026-10-07. Sources are listed at the end of the page.

What is a RAG agent?

Anthropic’s “Building effective agents” (19 December 2024) separates two designs. Workflows are “systems where LLMs and tools are orchestrated through predefined code paths.” Agents are “systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.” A standard RAG pipeline is a workflow. A RAG agent is the same retrieval placed under an agent’s control.

The LangGraph agentic RAG tutorial shows the shape. The model decides whether to call a retriever tool or answer directly. After retrieval, a grading step scores whether the documents are relevant. If they are not, the question is rewritten and retrieval is tried again. If they are, the model answers from them. The loop is the point, and the loop is also the risk.

How is this page different from the other agentic RAG pages?

PageWhat it covers
Agentic RAGThe pattern, how it differs from a fixed pipeline, its architecture and cost per query.
RAG vs MCPWhen to retrieve from an index and when to give an agent live tool access.
RAG pipelineThe stages of the fixed pipeline that an agent’s retrieval tool still depends on.
This pageWhat goes wrong when an agent controls retrieval, and the controls that bound it, with Swfte’s status for each.

What extra failure modes does an agent add to RAG?

Failure modeWhat happensControl
LoopThe grader keeps rejecting results, the question is rewritten again and again, and nothing converges.A maximum number of iterations. Anthropic recommends stopping conditions such as this.
Runaway costEvery rewrite, retrieval and check is another model call. Anthropic reports that agents typically use about 4 times more tokens than chat and multi-agent systems about 15 times more.A token or spend budget per run and per workspace, with an alert before the cap.
Tool misuseThe agent calls a tool with the wrong arguments, or one with side effects it should not have had.Least privilege for each tool, typed inputs and approval before any write.
InjectionRetrieved text carries instructions. LangChain’s documentation says no prompt or delimiter strategy fully prevents this.Treat retrieved text as data, restrict the agent’s tools and validate outputs. OWASP lists human approval for sensitive operations.
Compounding errorsA weak first retrieval shapes every later step. Anthropic notes that agents bring a potential for compounding errors.Grade retrieval at each step, log it, and stop for a person when confidence stays low.
Permission wideningThe agent searches with a broad service identity and returns what the asking person could not open.Retrieve as the person the agent acts for, filtered inside the query.

Which controls should a RAG agent have before it goes live?

  1. Step limit

    Set a hard maximum of retrieval and rewrite steps. When it is reached the agent answers with what it has, says retrieval was incomplete, or hands over to a person.

  2. Budget

    Set a token or spend ceiling for each run and for the workspace. Alert before the ceiling. A limit that only logs after the money is gone is not a control.

  3. Approval gates

    Require a person to approve any action that changes something outside the agent: sending, writing, paying, deleting. LangGraph’s interrupt pattern pauses a run at such a point and resumes on the person’s decision. See human in the loop.

  4. Trace logging

    Record each step: the query as written and as rewritten, the filter, the chunks returned, the grade, the tool calls and the final answer. Without the trace you cannot tell a retrieval problem from a model problem.

  5. Sandboxed tests

    Anthropic recommends extensive testing in sandboxed environments along with guardrails. Test with hostile documents in the corpus before real users arrive.

When should you not build a RAG agent?

Anthropic says agents suit open-ended problems where the number of steps cannot be predicted, and that their autonomy means higher costs. If your questions are lookups against one corpus, a fixed pipeline is cheaper, easier to test and has no loop to bound. Add the agent when a measured set of failures needs it: questions that span sources, or that need a second search after the first.

Where Swfte fits: what is built for a RAG agent

The existing page agents on the brain holds the full status table. In short:

PartStatusNote
Graph of people, groups and reporting lines the agent can readBuiltThrough the appliance’s local API, scoped by token.
Permission-filtered search over documentsIn progressLibraries exist. Endpoints are not finished.
Context-package API and MCP server for agentsRoadmapDesigned for. Not built.
Agents built in Studio; Nexus capturing agent actions, enforcing policy and tracing themBuiltNot yet wired to the company brain.
Spend and token caps in ConnectBuiltWorkspace and per-model caps, with alerts and an action when a cap is hit, such as downgrading to a cheaper model. They bound cost. They do not count steps.
Human approval before an actionBuiltSeveral approval mechanisms exist in separate parts of the platform. See human in the loop.

A step limit belongs in the agent definition or the framework you run. This page does not state that Studio exposes one. If your agent is a single fixed pipeline over one corpus, you do not need Swfte for this: a framework such as LangGraph and your own limits will do. Governance policy is described on platform governance.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is a RAG agent?

A RAG agent is a language model that controls its own retrieval. It decides whether to search, writes and rewrites the query, grades what comes back and loops until the evidence is good enough or a limit is reached. A standard RAG pipeline runs the same fixed steps for every question.

What is the difference between RAG and a RAG agent?

In RAG, code fixes the steps: retrieve, then generate. In a RAG agent the model chooses the steps, which Anthropic describes as dynamically directing its own process and tool use. That helps with multi-step questions and costs more, with a risk of loops and compounding errors.

How do you stop a RAG agent from looping?

Set a maximum number of iterations and decide in advance what happens at the limit: answer with partial evidence, say retrieval was incomplete, or escalate to a person. Add a spend budget as a second stop. Log every rewrite so you can see why the loop did not converge.

Does a RAG agent make prompt injection worse?

It gives injection more to work with. An agent reads more retrieved text and holds tools that can act on it. LangChain’s documentation says no prompt or delimiter strategy fully prevents indirect injection, so limit the agent’s tools, retrieve as the asking person and require approval before any write.

Is a RAG agent more expensive than RAG?

Usually, per question, because each rewrite, retrieval and check is another model call. Anthropic reports that agents typically use about 4 times more tokens than chat and multi-agent systems about 15 times more. That figure is for its own research system. Measure your own agent’s cost per answered question.

Can Swfte agents read the company brain today?

Not through a finished agent interface. The graph and its local API are built. Permission-filtered document search is in progress. The context-package API and MCP server are on the roadmap, as is wiring Studio and Nexus into the brain. See the agents on the brain page for the full table.

Show us the agent you want to run and we will map the controls it needs.

Centralise your knowledge in Cortex

The desktop AI workspace: 20+ providers, local models, knowledge bases with RAG that cite their sources, MCP tools and agents, with sensitive work staying on the laptop by default.