How Swfte builds with Cortex / Automated review

Automated review steps: what finds problems and what approves them

The difference between a step that finds problems and a step that approves a change, what Swfte runs today, and where our own review is still manual.

Automation is good at finding problems and poor at accepting responsibility for them. This page separates the two jobs. It lists the review steps Swfte runs today, including the ones that are weaker than they sound, and it says what Nexus gates do. Cortex does not front our own review. The practices are labelled in use, built, or designed for.

Finding and approving are different jobs

A review step does one of two things. It finds: it reads a change, runs a check, and reports what looks wrong. Or it approves: it takes responsibility for letting the change go. Automation can find tirelessly, at any hour, across everything you point it at. It cannot be accountable. An approval is a decision made by a named person who can be asked why, so an automated step should feed that decision and never replace it.

We keep the two apart on purpose. A finding is cheap to produce and cheap to dismiss, so we want plenty of them and we want them to carry evidence. An approval is expensive and rare, so we want it to be deliberate and recorded. When a step blurs the two, for example an AI reviewer whose pass counts as sign-off, we call it automated QA and we do not call it approval.

What Swfte runs today

Each of these is in use at Swfte. None of them is Cortex.

  • Adversarial, read-only review documents

    We write review documents that attack a system and change nothing. Each finding is numbered, and when a fix is made the document gains a closure table that links each finding to the change that closed it. The documents do not name their reviewer. The wording points to our multi-agent harness, but we state that as an inference.

  • Completion-gate ledgers

    A ledger lists the gates for a task: the command to run, the result expected, and the evidence recorded when it was run. A box ticked without evidence counts as unmet. The stronger ledger is the one in our Cortex application repository, where each gate carries a command, an expected result and recorded evidence.

  • Build checks

    Scripts that run before every build fail it when a rule is broken: a missing entry in the slug lockfile, or a model price row past its allowed age. These are the most reliable review steps we have, because they run every time and cannot be talked out of a failure.

  • Security checks in CI

    The website runs secret scanning, static analysis and a dependency audit in continuous integration, and a weekly audit that gates on a minimum score. These look for leaked keys, risky code patterns and known-vulnerable packages. They find. A person still decides what to do with each result.

  • Retracting claims that cannot be backed

    Claims honesty is a review practice for us. We have withdrawn security-posture, reliability and customer-count claims that nothing supported, and labelled testimonials as illustrative. It is done by people and agents reading the copy, not by an automated check, because no claims linter exists.

The limits, stated plainly

Most entries in our website gate ledgers are manual evidence rather than runnable commands. Only a handful of gates carry a command that a machine can re-run. That is useful as a record of what a person checked, and it is weaker than it looks, because nothing re-verifies it later. Likewise, none of our hosted CI jobs run the website tests or the build checks. We run those locally.

Our content pipelines are defined as Swfte workflows with an AI fact-check reviewer and no human step. That is automated QA, not approval, and we have no evidence in the repository that they are running today. In our Cortex application repository, hosted CI stopped running for a period and we ran the same gates locally from one script. An AI pull-request review workflow exists there, but we have no posted review on record, so we claim nothing about it.

What Nexus gates do

Nexus puts two kinds of gate around a coding agent. The first is a policy gate on each tool call. It runs on the machine, in the middle of the agent's work, and answers allow, deny or ask. You can run it in an audit mode first and move to enforcement once you trust the rules. The second is the completion gate, run with the nexus gates command, which checks that the gates you wrote for a task have recorded evidence.

Inside Cortex, a chat assistant can call tools to start verification, read gate results and read the audit trail for a coding task. All of that is built in the product. We use completion gates ourselves, as described above. We do not claim that Nexus gates are the thing standing between a pull request and our main branch. A gate finds that a rule was broken. Who accepts the change is still a person.

How a gate ledger works

This is the practice we use for large tasks, and the one we suggest you copy first.

  1. 01

    Write the gate before the work

    State what must be true when the task is done, before anyone starts. Writing the gate afterwards invites a gate shaped to fit what happened. Each gate needs a name a stranger would understand and a single outcome that is either met or not.

  2. 02

    Give it a command and an expected result

    Where a machine can check the outcome, write the command and the result to expect. Where only a person can check it, write the evidence they must record. Prefer the command whenever one exists, because a command can be run again by someone else.

  3. 03

    Run it after a person approves the run

    Checks run only after a person has approved the command, so a gate cannot become a way of running arbitrary code. Record the output and the exit code beside the gate, not a summary of them, so a reader can see what was actually run.

  4. 04

    Re-verify what a helper reports

    When a helper agent reports a gate as met, the parent task checks the evidence itself before it marks the gate as done. A report is a claim. The recorded evidence is what counts, and it must be there to read.

  5. 05

    Treat a ticked box without evidence as unmet

    The rule that gives the ledger its teeth is simple: no recorded evidence, no pass. It stops a long task from drifting into a feeling of being finished, and it makes a half-done task visible to the next person who opens the ledger.

What to automate first

Start with checks that fail loudly and never need judgement: a build that stops when a rule is broken, a scan for leaked secrets, a test that must pass. They are cheap, they run every time, and they teach the team to trust a red result. Then add read-only review of changes, where an agent reads a diff and reports numbered findings with evidence. Only after that consider anything that looks like a decision.

Resist the temptation to start with an automated approver. The value of a review step is in the problems it finds early, and the value of an approval is that someone owns it. If you automate approval before you have good findings, you collect sign-offs that nobody would defend. If you have good findings and a person who reads them, you can decide later which low-risk approvals to delegate, under limits.

A worked example, labelled design intent

This is how we would review a pull request that changes an approval flow. It is design intent, not a description of a review we have run. A read-only agent reads the diff and produces numbered findings: a check that could be skipped, a path with no recorded decision, a test that cannot fail. Each finding names the file and the evidence behind it.

A person then triages the findings, closes the false ones with a reason, and asks for fixes to the rest. Each fix gets a line in a closure table. The gates run, including the build checks, and a second person approves the merge, because the change touches approvals. If we have a real example from our own work, it will replace this one: <real example of an AI-assisted review catching a defect - founder to fill>.

What stays with a person

A person decides which findings matter and which are noise. A person accepts a risk, or refuses to, and records why. A person approves the merge and the release. A person decides whether an automated step is trustworthy enough to rely on, and reads its failures when it is not. In our own practice, some decisions are reserved to the owner outright, and the page on approvals describes them.

Nothing here lets a reviewer, human or automated, approve its own work. The reviewer that wrote the code should not be the one that clears it. That separation is a design aim, and the platform does not yet enforce it as a rule, so for now it lives in how a team is organised.

Automated review: what is in use, built and designed

The first five rows are real at Swfte today. The last three show where we stop short of a claim.

What is in use at Swfte, built in the product and designed for: Automated review
PracticeStatusWhat we can point to
Adversarial read-only review documents with numbered findings and closure tablesIn use at SwfteDated review documents exist for our Cortex application. The reviewer is not named, and we describe them as AI-assisted only by inference.
Completion-gate ledgers with command, expected result and recorded evidenceIn use at SwfteThe Cortex ledger is the strong example. Most website ledger entries are manual evidence.
Prebuild checks for rules and data freshnessIn use at SwfteSlug lockfile and price-row age checks fail the build. They run locally, not in hosted CI.
Security CI with secret scanning, static analysis and a dependency auditIn use at SwfteRuns on the website repository.
Retracting claims that cannot be backedIn use at SwfteDone by reading the copy. No automated claims check exists.
Content pipelines defined as workflows with an AI fact-check reviewerBuilt in the productDefined in our repository with no human step. That is automated QA, not approval, and we do not claim they run today.
Nexus policy gates and completion gates on coding-agent workBuilt in the productAllow, deny or ask on tool calls, and a gate check for completion. Cortex can start verification and read gate results.
AI review of pull requests as a merge controlDesigned forA review workflow file exists, but no posted review is on record. We claim nothing about it running.

How to read the status

  • In use at Swfte. We can point to it in our own repositories, pipelines or commit history.
  • Built in the product. The product can do this today. We make no claim that we run it on ourselves.
  • Designed for. How you can do it. Design intent, not a statement about what we have done.

Frequently asked questions

Does Swfte run this on its own company information?

In part. We run review documents, gate ledgers, build checks and security CI on our own repositories. We do not run our review through Cortex today. Our content pipelines have an AI fact-check step with no human approval, which is automated QA. An AI pull-request review workflow exists, but we have no posted review to point to.

What is the difference between automated review and approval?

Review finds problems and reports them with evidence. Approval accepts responsibility for letting a change go, and a named person can be asked why. Automation is well suited to the first and unsuitable for the second. We call an AI reviewer with no human step automated QA, not approval.

What is a completion gate?

It is a condition for calling a task done, written before the work starts, with a command to run, the result expected and the evidence recorded. A ticked box without recorded evidence counts as unmet. It stops a long task from drifting into a feeling of being finished.

Are your gates all machine-checkable?

No. Most entries in our website gate ledgers are manual evidence rather than runnable commands. Our Cortex application ledger is stronger, with a command, expected result and recorded evidence for each gate. We would rather say that than imply a level of automation we do not have.

Do you use AI to review pull requests?

A pull-request review workflow exists in one of our repositories, but hosted CI stopped running there for a period and we have no posted review on record. So we do not claim it as a practice. We do use read-only AI-assisted review documents, which a person reads and triages.

What should we automate first?

Start with checks that fail loudly and need no judgement: a build that stops on a broken rule, secret scanning, tests that must pass. Then add read-only review that reports numbered findings with evidence. Leave approval with a person until you trust the findings and can delegate low-risk cases under limits.

What do Nexus gates do?

Nexus has a policy gate that answers allow, deny or ask on each tool call a coding agent makes, and a completion gate that checks the evidence for the gates you wrote. Both are built in the product. Neither replaces a person approving a merge.

Take automated review further with Swfte

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.

Centralise your knowledge in Cortex

The desktop AI workspace: 20+ providers, local models, knowledge bases with RAG that cite their sources, MCP tools and agents, with sensitive work staying on the laptop by default.