How Swfte builds with Cortex / Improvement loop

The improvement loop: observe, propose, approve, ship, measure

The five stages by which an organisation can improve what it runs with agents, what exists at each stage today, and which parts are still only a design.

The loop is observe, propose, approve, ship, measure, and it runs through everything else on these pages. This page takes each stage in turn and says what exists. The honest summary is that parts of every stage are built, a personal version runs for one person on one device, and a company-wide loop across code and operations is design intent. No measured outcome is published.

The loop, and where it honestly stands

The loop has five stages. You observe what is happening. An agent proposes a specific change. A person approves, edits or rejects it. The change ships through the normal gates. You then measure whether it did what you wanted, and what you learn becomes the next thing you observe. The point of the shape is that the agent proposes and a person decides at the one stage where judgement is needed.

Here is where it stands. Pieces of every stage are built in the product, and some of them we use. A personal loop is built in Cortex and watches one person. Missions, the structured way to run a multi-step piece of work, have one documented run on an isolated validation stack and are not deployed to production. A company-wide loop across code and operations is design intent. We publish no measured outcome, because we have none to publish.

Observe: three sources, at three different scales

The first source is the Nexus ledger. When we run coding agents through Nexus capture, it records what they did as events on the machine: the tool actions, the file changes and the outcome of each turn. That is in use at Swfte as a record, and it is the closest thing we have to an observation of our own agent work. The second is session and issue observability: events from products go to a service that groups them into issues and uses an AI classifier. It is built in the product.

The third is a personal loop inside Cortex. It watches one person's activity on their own device, finds recurring patterns, and keeps the results on that device. It is built, and it observes one person. It is not a view of a company, and we have not confirmed that it is switched on in the shipped build. People choose what is worth watching. An observation that nobody chose is noise with a good user interface.

Propose: turning an observation into something specific

A proposal has to be specific enough to approve or refuse. The triage command shows the idea in a small form. Given an issue, it returns a diagnosis with a likely cause, the component and suggested fix text. It does not open a pull request. The personal loop in Cortex goes further and gates its own ideas with four questions: is there evidence, is there a named action, is the cost measured, and is there a way to show it wrong.

If a candidate fails any one of those four questions it is dropped. That discipline is worth copying whatever tool you use: an opportunity without evidence, an action, a cost and a falsifier is a hunch. Missions use the same instinct at larger scale: a plan is priced before launch, so a person approves a cost and a shape rather than a promise. In every case the agent proposes and does not apply.

Approve: where a person decides

Approval is the stage the other four exist to serve. In the product it takes several forms: the approval cards that Nexus raises in a chat, the human input and gate steps in workflows, and the human steps in a mission, which has its own final acceptance before a piece of work counts as done. Each of them records who decided. They are separate mechanisms, and the page on approvals explains the differences.

Our own approval practice is smaller. Decisions about a live payment test and about the release signing key are kept by the owner, and a rule in our repositories tells AI assistants to propose before they execute. A person approves. Nothing in this loop lets an agent approve its own proposal, though the platform does not yet enforce that the approver differs from the requester. For now that separation is a team choice.

Ship: through the gates you already have

An approved change is not shipped by the loop. It is shipped by your normal release process, with its tests, its review and its gates. In our Cortex application a merge releases only if the version has been bumped and the gates are green. When hosted CI stopped running for a period, we ran the same gates locally from one script, and we recorded that this left nothing gating a merge in the usual way.

For larger tasks we also write completion gates before the work begins, with a command, an expected result and recorded evidence. These are in use at Swfte. The loop adds no new route to production. If it did, it would be a way round the controls the rest of this site describes. A change that the loop proposes and a person approves still has to pass what any change has to pass.

Measure: did it work, or is that still unproven

Measurement is the weakest stage, and we would rather say so. The personal loop in Cortex keeps a ledger of each opportunity as accepted, dismissed or built, and a hit rate follows from it. We have not verified that anyone's outcomes have been measured with it. Missions carry a five-state measure in which unproven is a separate state from missed, so a result you cannot yet judge is not counted as a failure.

That distinction is worth keeping. "We have not shown it worked" and "it did not work" are different findings, and mixing them up leads teams to abandon good changes or to claim wins they have not earned. People judge whether a change worked. They do it against the reason the change was made, written down before it shipped. We publish no measured outcome from any loop, ours or anyone else's.

Status by stage

What exists at each stage, with the label that applies.

  • Observe: Nexus ledger, observability, personal loop

    In use at Swfte: capture of our coding agents' actions through Nexus. Built in the product: session and issue observability with an AI classifier, and the personal on-device loop in Cortex that watches one person's activity. A view of a whole company is designed for.

  • Propose: triage, the four-question gate, mission plans

    Built in the product: a triage diagnosis with suggested fix text, a four-question gate that drops weak opportunities, and mission plans that are priced before launch. An agent that proposes a diff or a workflow change for approval is designed for.

  • Approve and ship: gates, cards, release checks

    Built in the product: approval cards, workflow gates and mission human steps, with the decider recorded. In use at Swfte: completion gates and owner-decided steps. Shipping still goes through your normal release process. A single approval inbox is designed for, not built.

  • Measure: ledgers and the five-state measure

    Built in the product: an accepted, dismissed and built ledger with a hit rate, and a five-state mission measure where unproven is not missed. Not published: any measured outcome, and we will not invent one. Placeholder: <measured outcome of the improvement - founder to fill>

A worked example, labelled design intent

Here is how we would run one turn of a company-wide loop. It is design intent, not a description of something we have done. Observation shows that a particular error recurs in one product screen. An agent turns that into a proposal: change this check, and here is the evidence. The proposal passes the four questions, so it is not a hunch. A named person reads the evidence and approves it with a note.

The change then goes through the normal tests, review and release. After release, the team compares the error rate with the reason the change was made, and records one of three findings: it worked, it did not, or it is still unproven. That finding is the next observation. If we have a real change of our own to show, it will replace this one: <real example of a shipped improvement from the loop - founder to fill>.

What stays with a person

People choose what to watch, because an agent that watches everything finds everything and proves little. People approve, because approval carries responsibility that an agent cannot hold. And people judge whether the change worked, because that judgement depends on the reason it was made, which is a human reason. The agents observe, draft and propose. They do not choose, accept or conclude.

Read the other pages for the parts: internal knowledge for what the loop can see, the engineering workflow for how changes are made, automated review for how problems are found, and approvals for how decisions are recorded. The evidence table below is the plain list of what is in use, what is built and what is still design intent for the loop as a whole.

Improvement loop: what is in use, built and designed

The loop is mostly built in pieces. Only some pieces are things Swfte does today, and a closed loop across the company is not one of them.

What is in use at Swfte, built in the product and designed for: Improvement loop
PracticeStatusWhat we can point to
Nexus capture of coding-agent actions as an observation recordIn use at SwfteMachine-local. We claim a record of agent actions and no more.
Completion gates written before large tasksIn use at SwfteGate ledgers exist for our work. Most website entries are manual evidence rather than runnable commands.
Session and issue observability with an AI classifierBuilt in the productGroups events from products into issues. We make no claim that we act on it as a company loop.
Personal on-device loop in CortexBuilt in the productObserves one person's activity and keeps an accepted, dismissed and built ledger with a hit rate. We have not confirmed it is switched on in the shipped build.
Triage diagnosis and the four-question opportunity gateBuilt in the productTriage returns suggested fix text and opens no pull request. The gate drops any candidate that fails a question.
Missions: plans priced before launch, human steps, five-state measureBuilt in the productOne documented run on an isolated validation stack. Not deployed to production, and not live for customers.
A company-wide closed loop across code and operationsDesigned forStrategy only. Nothing joins observation to measurement across the company today.
A published measured outcome from any loopDesigned forNone exists. We publish no figures until there is one we can stand behind.

How to read the status

  • In use at Swfte. We can point to it in our own repositories, pipelines or commit history.
  • Built in the product. The product can do this today. We make no claim that we run it on ourselves.
  • Designed for. How you can do it. Design intent, not a statement about what we have done.

The improvement loop

This page walks all five stations. Pieces of each stage are built, a personal loop runs for one person, and a closed company-wide loop is design intent.

  1. 01 · this page

    Observe

    Collect what is happening: errors, struggle signals, review findings, cost, support questions, usage.

    People choose what is worth watching.

  2. 02 · this page

    Propose

    An agent turns an observation into a specific proposal: a diff, a workflow change, a policy edit.

    The agent proposes. It does not apply.

  3. 03 · this page

    Approve

    A named person reviews the proposal and its evidence, and approves, edits or rejects it.

    A person decides, and the decision is recorded.

  4. 04 · this page

    Ship

    The approved change goes out through the normal gates: tests, review, release.

    The release gates still apply.

  5. 05 · this page

    Measure

    Compare the outcome with the reason you made the change, then feed it back into what you observe.

    People judge whether it worked.

The last step feeds the first. What you measure becomes the next thing you observe.

Frequently asked questions

Does Swfte run this on its own company information?

Only in pieces. We capture our coding agents' actions through Nexus, write completion gates and keep some decisions with the owner. We do not run a closed loop across our code and operations, and the personal loop in Cortex observes one person. A company-wide loop is design intent, and we publish no measured outcome.

What are the five stages of the improvement loop?

Observe what is happening, propose a specific change, approve it by a named person, ship it through the normal gates, and measure whether it did what you wanted. What you learn becomes the next observation. The agent proposes and a person decides at the approval stage.

What is the four-question gate?

It is a test an opportunity must pass: is there evidence, is there a named action, is the cost measured, and is there a falsifier, meaning a way to show the idea wrong. Failing any one drops it. The personal loop in Cortex applies it, and the discipline is worth copying with any tool.

Are Missions live for customers?

No. Missions are built, with plans priced before launch, human gates and final acceptance, but they have one documented run on an isolated validation stack and are not deployed to production. Treat them as built in the product and not yet available to rely on.

Does the loop change production on its own?

It should not, and the design does not allow it. An agent proposes, a person approves, and the change ships through your normal tests, review and release gates. The loop adds no new route to production. The platform does not yet enforce that the approver differs from the requester, so that is a team choice.

What does unproven mean in the mission measure?

Missions carry a five-state measure, and unproven is a separate state from missed. It means you have not yet shown whether the change worked. Keeping them apart stops a team from counting an unjudged result as a failure, or from claiming a win it has not earned.

Can I see a measured improvement from the loop?

Not from us. We have no measured outcome to publish, and we will not invent one. The personal loop keeps a ledger and a hit rate, and Missions have a five-state measure, but we have not shown results from them. Measure it against your own reason for the change.

Take improvement loop further with Swfte

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.

Centralise your knowledge in Cortex

The desktop AI workspace: 20+ providers, local models, knowledge bases with RAG that cite their sources, MCP tools and agents, with sensitive work staying on the laptop by default.