← The journal
Guides

Governed Agents: The Checklist Before Production

An eleven-item checklist for putting an AI agent into production under control, from identity to exit plan.

Swfte Journal / Guides

An AI agent is ready for production when you can say who it is, who answers for it, what it may touch, what it may do without asking, what is recorded, and how you stop it, and when you can check each of those rather than just state it. The checklist below has eleven items. Each has a reason, a way to check it, and a plain label for whether the Swfte platform builds it today or whether it is design intent that you cover with process. If a box cannot be ticked with evidence, the agent is not ready, whatever the demo looked like. The wider idea is on the governed agents page, and the principle behind it is that governance should run at runtime, not on paper, as argued in why AI governance must run at runtime.

What does production-ready mean for an agent?

It means the agent's authority is bounded and visible. A person reading the record can tell what the agent was allowed to do, what it did, who agreed, and what came of it. That is a higher bar than "it works". A tool that answers questions can be wrong and get corrected. An agent that sends, pays, deletes or changes records can do harm before anyone looks. So the bar rises with the agent's reach.

A rule of thumb: the checklist is cheap when run before launch and expensive when run after an incident. Do it before.

Who is the agent, and who answers for it?

1. Does the agent have its own identity and scoped credentials?

Why: if an agent borrows a person's login or a shared key, every action looks like that person, and revoking access means revoking theirs. How to check: list the credentials the agent uses and confirm each one belongs to the agent alone, with the narrowest permissions that work. Rotate one and confirm only this agent breaks. Status: this is a practice you apply. The platform does not hold an agent record that carries owner, risk level and approver today, so keep the list of credentials in your own register.

2. Is there a named owner?

Why: an incident needs someone to call, and an auditor needs someone to question. How to check: the owner is a named person with authority over the process, and they have accepted the role in writing. A team alias fails the check. Status: by process. Put the owner in the record you keep in item 4.

3. Is the scope of data and tools written down and enforced?

Why: over-broad access is the most common design flaw in a new agent. How to check: list the systems and data the agent may read, and the tools it may call. Then try something outside the list and confirm it is refused. The platform's policy engine has control points for the model call, the tool call, workflow nodes, outbound traffic, secrets and deployment, and policies can apply at platform, organisation, workspace or entity scope. Status: built, with a condition: enforcement is live only for runs that have a policy attached, so confirm yours does. See the platform's agent layer.

What is written down about the agent?

4. Do you keep a Trust Profile as a record?

Why: one page that states what the agent is, what it may do and who decides is the thing a reviewer reads first. The Swfte platform describes it as identity, owner, risk level, approved models, data classification, data residency, permitted systems, allowed and restricted actions, the human approval rule, retention, audit and policy set. How to check: the record exists, is dated, names the owner, and matches what the agent is actually configured to do. Compare them line by line. Status: the Trust Profile is a concept the platform describes; it is not yet a single stored object in the product. Today its parts sit in separate controls, such as policy, approvals and the run ledger. Keep the record yourself, in version control if you can, and treat a mismatch between page and configuration as a defect. See Trust Profile.

5. Do your policy verbs map to what the engine really does?

Why: governance language is often wider than enforcement. The framework talks about allow, deny, warn, filter, escalate and require human approval. The engine's verdicts are allow, redact, ask and deny. How to check: map each verb you rely on to a verdict you have tested. Redact covers filtering. Ask covers escalating to a person and requiring approval. There is no separate warn verdict, so if you want a warning, build it as monitoring. The engine also applies stricter outcomes first: a later guard can never soften an earlier one. Status: built. Write your mapping down so nobody assumes a verb exists that does not.

How are people kept in the loop?

6. Are approvals and timeouts defined?

Why: a gate with no timeout rule either blocks forever or, worse, gets worked around. How to check: for each action that needs a person, confirm who is asked, what happens if they do not answer, and that the answer is deny or wait, never approve. Run it. High-risk actions such as send, pay, delete, push and shell should always need a person. Do not count a Worker's approval mode as a human gate: Workers are headless and advisory, and that setting removes tools at deploy time rather than asking anyone. Status: built, in several separate parts of the platform rather than one shared model. Approver separation, roles and delegation are design intent. See approval workflows auditors can work with.

7. Is there a run ledger and a trace?

Why: without a record you cannot investigate, defend or improve. How to check: run the agent, read the events back, and run the verify call on the chain. The ledger is an append-only, hash-chained list per run covering policy decisions, approvals and gates. Confirm you can answer who, what, when and what next from it alone. Status: built. The ledger is readable as JSON and has no dedicated export yet, and an opt-in seal of the chain head exists, so ask whether it is on. Observability is covered in AI observability for agents and workflows.

Can you stop it, limit it and watch it?

8. Is there a way to stop it?

Why: the day you need it is the worst day to discover it does not work. How to check: rehearse it. Pause or stop a running agent and confirm what happens to work in flight. We can say this much: denying a workflow gate cancels the run unless a deny branch is wired. Anything beyond that, such as a stop control per agent or per tool, you should test in your own deployment before you rely on it. Status: partly built; verify what you depend on.

9. Does it start at L1 Assist or L2 Approve?

Why: trust is earned from evidence, not assumed. How to check: the first release recommends or drafts, and a person approves before anything is sent or changed. Write the promotion criteria first and the evidence you will collect. The levels are L1 Assist, L2 Approve, L3 Supervise, L4 Autonomous and L5 Adaptive. Status: designed for. The levels are a rollout language, not a field the platform stores or enforces today, so hold the current level in your own record. The full plan is in controlled autonomy: a practical rollout and on the autonomy page.

10. Are monitoring and an incident path in place?

Why: agents drift, inputs change and mistakes surface late. How to check: name who watches which signals, how an alert reaches the owner, and who may pause the agent out of hours. Run a tabletop exercise with a made-up bad output. The AI incident response runbook is a starting structure. Status: by process, supported by the ledger and traces.

11. Is there an exit plan?

Why: you may need to retire the agent, change the model or leave a vendor. How to check: confirm you can extract the agent's configuration, its knowledge sources, and the run records you must keep. Since the ledger has no dedicated export, test extraction before you need it. Decide what happens to in-flight work and to data the agent created. Status: by process.

What does the whole list look like at a glance?

ItemCheck in one lineStatus
Own identity and credentialsRotate one credential, only this agent breaksPractice you apply
Named ownerA person, accepted in writingProcess
Scope of data and toolsOut-of-scope call is refusedBuilt, needs a policy on the run
Trust Profile recordPage matches configurationDesigned, parts built separately
Policy verbs mappedEach verb tied to a tested verdictBuilt
Approvals and timeoutsUnanswered gate denies or waitsBuilt, in separate parts
Run ledger and traceVerify call passesBuilt
A way to stop itRehearsed, in-flight work knownPartly built
Start at L1 or L2Criteria written before launchDesigned
Monitoring and incident pathTabletop completedProcess
Exit planExtraction testedProcess

Starter shapes for several of these are planned on the governed agent templates page, and the zero trust approach to AI covers the access side in more depth.

What is built, and what is design intent?

Built in the product: a policy engine with allow, redact, ask and deny verdicts, approvals in several parts of the platform with timeouts that deny rather than approve, and a hash-chained run ledger with a verify call. Designed for: the Trust Profile as one stored record, autonomy levels as a setting, platform-enforced approver separation, approver roles, and a retention policy object. In use at Swfte: rules that our coding agents must follow, enforced by build checks, and dated review documents. Cortex does not front our own website work today. This checklist supports evidence for what you must show about an agent. It does not decide whether any system meets a regulation, which depends on your use case, jurisdiction, deployment and configuration, and it is not legal advice.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.