← The journal
Guides

Agentic Workflow: What It Is and How to Design One

An agentic workflow mixes fixed steps with model-driven decisions. Patterns, failure handling, approvals and cost.

Swfte Journal / Guides

An agentic workflow is a process in which one or more language-model steps decide something (what to retrieve, which tool to call, whether a draft is good enough) inside a structure you control. The structure, the order of steps and the limits are written by you. The model fills in the judgement. That is different from a fully autonomous agent, where the model also chooses the order of work. Last verified 2026-10-07.

This guide defines the terms from primary sources, sets out the five common patterns with a table of when to use each and what can go wrong, and then covers state, tools, failure handling, approvals, evaluation, cost and observability.

What is the difference between a workflow and an agent?

Anthropic's engineering post "Building effective agents", published on 19 December 2024, draws the line most teams now use. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths". Agents are "systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks."

The same post states the trade-off: "Workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale." It adds that agentic systems "often trade latency and cost for better task performance", and that you should consider when the trade-off makes sense.

Microsoft's Agent Framework documentation gives a similar rule in a table. Use an agent when the task is open-ended or conversational and needs autonomous tool use and planning. Use a workflow when the process has well-defined steps and you need explicit control over execution order. It then adds one sentence worth keeping on a wall: "If you can write a function to handle the task, do that instead of using an AI agent."

When does a fixed workflow beat a free agent?

Choose the fixed workflow when you can write the steps down. Typical cases are invoice intake, ticket triage, document review against a checklist and report generation. You get predictable cost, a clear place to insert a human check, and failures that point to one step.

Choose more agency when you cannot predict the steps in advance, such as an open-ended research task or a fault investigation where each finding changes the next question. Even then, Anthropic's advice is to add complexity "only when it demonstrably improves outcomes", to start with simple prompts and to add multi-step agentic systems only when simpler solutions fall short. In practice that gives an order of attempts: one well-prompted model call, then a fixed chain, then routing and parallel steps, then an orchestrator, and only then a free-running agent with tools.

What are the five patterns?

Anthropic names five workflow patterns built on an "augmented LLM", which is a model with retrieval, tools and memory. The "use when" column below quotes or closely follows its description. The risk column is this post's own reading, not Anthropic's.

PatternUse whenTypical risk
Prompt chainingThe task "can be easily and cleanly decomposed into fixed subtasks"; you trade latency for higher accuracyAn early mistake carries through every later step unless a check sits between them
Routing"There are distinct categories that are better handled separately, and where classification can be handled accurately"A wrong classification sends the case down the wrong path; unknown categories need a fallback
Parallelisation (sectioning or voting)Subtasks can run in parallel for speed, or "multiple perspectives or attempts are needed"Cost multiplies with every branch; merging conflicting outputs is its own problem
Orchestrator-workersYou "can't predict the subtasks needed", so a central model plans and delegatesThe planner can loop, over-delegate or lose the thread; needs hard limits
Evaluator-optimiserThere are "clear evaluation criteria" and iterative refinement "provides measurable value"The loop can polish the wrong thing, or never converge, without a stop rule

Most production designs combine two or three. A common one: route the request, run a fixed chain for each route, and add an evaluator step before anything is sent outside the organisation.

How should you handle state and memory?

State is everything the process needs to carry from one step to the next: the input, intermediate results, tool outputs, decisions and who approved what. Decide three things.

  • What is stored between steps. Keep it small and typed. Passing whole transcripts forward raises cost and spreads errors.
  • Where it lives. A durable store lets a run resume after a failure or a long wait for a person. LangGraph's documentation describes agents that "persist through failures" and run "for extended periods, resuming from where they left off". Microsoft's framework lists checkpointing and human-in-the-loop control, and the OpenAI Agents SDK has sessions as a persistent memory layer.
  • What is remembered across runs. Short-term working state and long-term memory are separate. Long-term memory needs the same retention and deletion rules as any other store of user data.

How do tools fit in?

A tool is a function the model can call: search, read a record, create a ticket, send a message. Anthropic's advice is to invest as much effort in the interface the model uses as you would in a human-facing one, with clear documentation and testing. Give each tool a narrow purpose, validate its inputs, return errors the model can act on, and separate read tools from write tools. Write tools are where approvals belong.

The Model Context Protocol (MCP) is a common way to expose tools. The OpenAI Agents SDK documentation shows per-tool approval settings for local MCP servers, for example requiring approval always for a delete tool and never for a read tool, and static or dynamic filters on which tools an agent may see. Treat that as a pattern to copy whatever framework you use.

How should failures be handled?

Assume every step can fail, time out, return something malformed or be wrong.

  1. Set stopping conditions. Anthropic notes it is common to include limits "such as a maximum number of iterations" to maintain control. Add a time limit and a spend limit too.
  2. Validate between steps. Check format and obvious sense before the next step consumes the output.
  3. Retry only what is safe to repeat. Reads are safe. A payment or an email is not, unless the step is idempotent.
  4. Have a fallback path. A smaller model, a different provider or a hand-off to a person.
  5. Test in isolation. Anthropic recommends "extensive testing in sandboxed environments, along with the appropriate guardrails."

Where do human approval gates belong?

Put a gate before any action that is hard to reverse, costs money, sends something outside the organisation or touches personal data. Anthropic's description of autonomous agents includes pausing "for human feedback at checkpoints or when encountering blockers". Design the gate so that the reviewer sees the exact action proposed, the evidence behind it and the alternatives, and so that a timeout means "denied", not "approved".

The aim is to remove gates one at a time as evidence builds, not to start without them. Swfte's controlled autonomy model uses that idea: five levels from Assist to Adaptive, with the level raised per action as a record of accepted work builds up. See human-in-the-loop AI for gate design in more detail.

How do you evaluate an agentic workflow?

Anthropic puts it simply: "The key to success, as with any LLM features, is measuring performance and iterating on implementations." Evaluate each step as well as the whole run. Score the router's classifications, the quality of each chain step, tool-call correctness and the final outcome, and track cost and latency next to quality. Keep a regression set of real cases and run it on every change to a prompt, a tool or a model. The method is set out in LLM evaluation.

How do you control cost?

Agentic systems trade cost for performance, so budget them. Set a maximum number of model calls and tokens per run, a per-user or per-workflow daily cap, and an alert when a run exceeds its norm. Use a cheaper model for routing and classification and a stronger one only where the evaluation shows it is needed. Cache results of tool calls that are safe to reuse. Measure cost per completed task, not per call.

What should you observe?

Record a trace for every run: each model call with its prompt version, each tool call with its arguments and result, each branch taken, each approval and who gave it, and the cost and latency of each step. The OpenAI Agents SDK lists tracing among its core primitives, and LangGraph's documentation points to deep visibility into agent behaviour through LangSmith. Whatever tool you use, you need to be able to answer "why did it do that?" for any run. See multi-agent workflow for how this changes when several agents cooperate.

Where does Swfte fit?

Swfte Studio is a builder for this kind of process, and you can use other tools for the same job.

  • Workflows with agent steps and human-input nodes. A Studio workflow can include a human-input node that pauses the run, assigns the task to a person, and continues on an approve or reject branch; a default timeout applies. Status: Built. See workflows and how to build an AI agent.
  • Policy and approvals. The platform has a policy engine with allow, redact, ask and deny outcomes, and several approval mechanisms in different products. Enforcement applies to runs that have a policy attached, and the approval mechanisms are separate rather than one shared inbox. Status: Built, with those limits.
  • Nexus wraps Claude Code and Codex with a policy gate (allow, deny, ask), an audit trail and completion gates. Status: Built.
  • Autonomy levels. The five-level model is a design framework. L1 to L3 describe how approval and monitoring change; L4 and L5 are designed for and are not offered as settings today.
  • Swfte does not publish a benchmark or an uptime figure for agent runs.

If your process has five steps and one model call, you do not need a platform. Write the function.

Sources and last verified

Last verified 2026-10-07. Every dated or technical fact in this post was read from the pages below on that date. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is an agentic workflow?

An agentic workflow is a process where a language model makes some of the decisions, such as what to retrieve, which tool to call or whether a draft is good enough, inside a structure you define. You control the order of steps and the limits, and the model supplies judgement within them. A fully autonomous agent differs because the model also chooses the order of work.

What is the difference between a workflow and an agent?

Anthropic defines workflows as systems where models and tools are orchestrated through predefined code paths, and agents as systems where models dynamically direct their own processes and tool usage. Workflows offer predictability for well-defined tasks. Agents suit cases where flexibility and model-driven decisions are needed, at a cost in latency and spend that you should weigh.

When should you use a fixed workflow instead of an agent?

Use a fixed workflow when you can write the steps down. Microsoft advises that if you can write a function to handle the task, you should do that instead of using an AI agent. Anthropic advises adding complexity only when it demonstrably improves outcomes. Fixed steps give predictable cost, an obvious place for a human check, and failures that point to one step.

What are the five agentic workflow patterns?

Anthropic names prompt chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser, all built on an augmented model with retrieval, tools and memory. Chaining splits a task into fixed steps, routing sends inputs to specialised paths, parallelisation runs branches at once, an orchestrator delegates unpredictable subtasks, and an evaluator loop refines output against clear criteria.

Where should human approval go in an agentic workflow?

Put an approval before any action that is hard to reverse, costs money, sends something outside the organisation or touches personal data. Show the reviewer the exact action, the evidence and the alternatives, and make a timeout mean denied. Start with more gates and remove them one at a time as a record of accepted work builds up for that action.

How do you stop an agentic workflow from running up costs?

Set limits before the first run: a maximum number of model calls and tokens per run, a time limit, a daily cap per workflow and an alert when a run exceeds its norm. Anthropic notes that stopping conditions such as a maximum number of iterations are common. Use a cheaper model for routing, and measure cost per completed task rather than per call.

Can Swfte Studio build agentic workflows?

Yes, within the limits stated here. Studio workflows can include agent steps and human-input nodes that pause a run, assign a person and continue on approve or reject, which is Built. Policy enforcement applies to runs that have a policy attached. The L4 and L5 autonomy levels are designed for, not offered as settings today. A simple process may not need a platform at all.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.