← The journal
Strategy

Controlled Autonomy: A Practical Rollout Plan from Assist to Adaptive

A rollout plan for controlled autonomy: move AI agents from L1 Assist to L5 Adaptive on evidence.

Swfte Journal / Strategy

Controlled autonomy means an AI agent is given more independence in small steps, each one justified by evidence from the step before, and each one taken back if the evidence turns. It is not a switch between "manual" and "autonomous". The platform describes five levels: L1 Assist, L2 Approve, L3 Supervise, L4 Autonomous and L5 Adaptive. This post gives a practical plan for moving an agent up that ladder: what each level means, what you must have in place before promoting, what evidence to collect, what should trigger a demotion, and a phased plan that avoids made-up targets. The ideas sit inside the wider Sovereign Intelligence framework, and the reference page on the platform is controlled autonomy.

What are the five levels?

The levels below use the platform's own definitions. Autonomy is set per agent and per action, so one agent can sit at different levels for different actions.

LevelNameWhat it means
L1AssistAI recommends. A person does the work and makes every decision.
L2ApproveA human approves before anything happens.
L3SuperviseAI acts within limits and is monitored.
L4AutonomousAI works independently within strict policy and risk bounds.
L5AdaptiveAI improves within controlled boundaries.

Two clarifications prevent common confusion. First, a higher level is not a better level. Many agents should stay at L2 or L3 permanently, because the right level depends on risk, reversibility and the evidence you hold. Second, the levels change four things at once: the approval rules, the monitoring, the risk boundaries and the depth of audit. Promoting an agent without changing those four is just flipping a flag.

Why not go straight to full autonomy?

Because capability without control is not enterprise-ready, and trust in an agent is a property you observe, not one you declare. An agent that looks excellent in a demo may fail on the tenth unusual input. The cost of that failure scales with what the agent can do, so the rational policy is to grant authority at the pace that evidence arrives.

There is also a regulatory angle. For systems that fall under high-risk rules in the EU AI Act, human oversight is a named requirement (Article 14), and record-keeping of events is another (Article 12). The timing of those obligations has shifted: the Digital Omnibus on AI deferred stand-alone high-risk obligations to 2 December 2027, per Gibson Dunn's summary. Whether and how any of it applies to your system is a question for counsel. Designing with oversight and records from the start costs little. Retrofitting them later costs a lot.

What must be in place before the first promotion?

Do not promote an agent past L2 until the basics exist. Think of these as the entry criteria for the whole ladder.

  • A Trust Profile. One record per agent: identity, owner, risk level, approved models, data classification, data residency, permitted systems, allowed actions, restricted actions, human approval rule, retention, audit and policy set. See Trust Profile.
  • Its own identity and scoped credentials. No shared keys. See step 5 of the build guide.
  • Runtime enforcement. Policy must be able to Allow, Deny, Warn, Filter, Escalate or Require human approval on real calls, as discussed in why governance must run at runtime.
  • A trace for every action. You cannot promote on evidence you do not have.
  • A named owner and an incident path. Someone is accountable, and there is a known way to pause the agent.
  • A way to stop it. A tested kill switch or pause, at agent level and at tool level.

If one of these is missing, the answer to "can we promote?" is no, whatever the agent's performance.

How should you decide on promotion?

Define promotion criteria before you start, in writing, so the decision is not made under pressure. The criteria should be about evidence types, not invented thresholds. Where you need numbers, set them with your owners, such as <acceptance threshold - set by agent owner>, rather than borrowing a figure from a blog.

PromotionEvidence to collectQuestions to answer
L1 to L2Acceptance and edit rates of recommendations, cases where the human overrode the agent and whyAre recommendations usually right, and are the wrong ones caught?
L2 to L3Approval history, rejected drafts, issues the agent caught, time the approver spends per itemDo approvers mostly accept without edits for a defined class of action?
L3 to L4Monitored runs, limit hits, escalations, exceptions, rollback testsAre limits tight, are exceptions rare and understood, does rollback work?
L4 to L5Drift and regression results, change history, test results on proposed behaviour changesCan changes be tested, bounded and reversed before they reach production?

The record from each level is the input to the next, which is the closed intelligence loop applied to autonomy itself.

What does the AI Procurement Agent look like at each level?

The worked example used across the platform pages makes this concrete. The AI Procurement Agent can read approved supplier information, analyse contracts, compare pricing, prepare purchase recommendations and create draft purchase orders. It cannot access unrelated employee data, approve its own high-value transaction, make payments or modify restricted records. Purchases above a threshold, contractual changes and sensitive external communications require approval. It records identity, data accessed, model used, output, tools called, policy applied, decision, approval, action and outcome.

  • At L1 Assist, the agent reads approved supplier information, compares pricing and recommends a purchase. A buyer does everything else. Track how often buyers accept the recommendation.
  • At L2 Approve, it prepares draft purchase orders and a buyer approves each one before it goes anywhere. Track edits, rejections and contract issues it caught.
  • At L3 Supervise, it creates draft purchase orders within limits and routes them on. Purchases above threshold, contractual changes and sensitive external communications still require approval. It still cannot make payments. It is monitored live, and limits are tightened or relaxed based on the record.

Notice that "cannot make payments" holds at every level in this example. Some restrictions are permanent. Autonomy grows inside the boundary of what the agent may ever do, and the boundary itself is changed only by the accountable owner, recorded in the Trust Profile. An agent should never be able to raise its own level. You can read more about the agent model on the governed agents page.

What are good demotion triggers?

Promotion rules without demotion rules are one-way doors. Write the reverse path in advance.

  1. Policy breach. The agent attempts a Denied action, or a Deny it triggered was enforced late. Drop a level pending review.
  2. Drift. Acceptance or edit rates move away from the baseline in a way nobody expected, or the model version changed without a re-evaluation.
  3. Data change. A new data source or a change in data class means the earlier evidence no longer applies.
  4. Tool change. A new tool, a changed permission or a changed downstream system resets evidence for the actions that touch it.
  5. Incident. Any harm or near-miss with a plausible link to the agent. Pause first, investigate second.
  6. Ownership gap. The named owner leaves or the team is reorganised. An agent with no accountable owner should not run above L2.
  7. Evidence gap. Traces are missing for a period. Treat it as a failure of the control, not a reason to assume things went fine.

Demotion should be fast, cheap and blameless. If demoting an agent is a political event, teams will hide problems.

What is a practical phased plan?

The plan below uses phases rather than durations or target figures, because the right pace depends on the agent's risk and on how quickly evidence arrives. If you need calendar dates, set them with your owners, for example <phase length - set per agent>.

Phase 0: Foundation. Choose one agent and one narrow process. Complete the Trust Profile, identity, enforcement and tracing. Run the AI estate inventory so the agent is not the only thing you know about. Define promotion and demotion criteria.

Phase 1: Assist. Run the agent at L1 on real work with real users. Record every recommendation and the human's response. Fix the cases it gets wrong, usually retrieval gaps and missing context, before touching autonomy. Review the traces for anything that would have been Denied.

Phase 2: Approve. Let the agent prepare actions in draft or staging. A human approves each one. This is where most of the evidence is generated, and where approver fatigue shows up. If approvers click through without reading, the control is decorative, so sample and review their decisions too.

Phase 3: Supervise. For a defined class of low-risk, reversible actions, let the agent act within explicit limits. Everything outside the limits still routes to a person. Monitor live, with alerts on limit hits and exceptions. Test rollback on purpose.

Phase 4: Autonomous, selectively. Only for processes where the limits, rollback and monitoring are proven and the process is well understood. Policy-defined gates replace per-item approval, with sampled human review of the trace. Many agents will not and should not reach this phase.

Phase 5: Adaptive, rarely. Allow an agent to improve its own behaviour, such as refining prompts, retrieval or routing, only within boundaries that it cannot modify, with each change tested, versioned and reversible. Change gates and drift monitoring are the controls here.

Across all phases, repeat the same loop: collect evidence, review it with the owner, decide, record the decision in the Trust Profile.

How do you keep humans meaningfully in the loop?

Human oversight fails in boring ways. Approvers are overloaded, shown too little context, or rewarded for speed. Four practices help.

  • Show reasons, not just actions. The approver should see the data used, the policy that triggered approval and the agent's proposed action, not a bare yes or no.
  • Route to the right person. Escalation to someone without authority is just delay.
  • Measure the approvers. Track how often approvals are changed on review and how long they take.
  • Reduce approvals by narrowing, not by skipping. If approval volume is too high, tighten the class of action that needs it, based on evidence, rather than removing the control.

These are also the first things a reviewer will ask about. The CISO page and the Chief AI Officer page list the vendor questions that probe them, and the operational sovereignty deep dive explains why decisions about what AI may do on your behalf should stay with you.

How does Swfte support this?

The platform is built so that approval steps, limits and records are part of how agents and workflows run. Studio builds agents and workflows with approval steps, and Nexus captures actions, enforces policy in flight and traces agents. Higher autonomy levels are the platform's position on where governed agents can go, and each deployment sets the levels it allows. The platform provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. See the trust centre for what is true today. For the build sequence, step 8 of the guide covers autonomy levels in more detail.

Where to go next

Pick one agent and write its promotion and demotion criteria this week, before anyone asks for more autonomy. Then check your overall footing with the readiness self-assessment, which runs in your browser and sends nothing anywhere. If you want to work through a specific agent with us, talk to our team. The glossary defines the terms, and the Sovereign Intelligence pillar places autonomy in the full picture.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.