From Company Brain to Custom Model: A Practical Path
Ten steps from a company brain to a custom domain model, with what Swfte provides today at each stage.
The practical path from a company brain to a custom model runs in ten steps, and most of the value arrives before any training happens. You connect identity and structure first, add documents and systems, use the brain for retrieval and agents, and only then look for the one task where the model's behaviour, rather than its knowledge, is what holds you back. That task is where a custom model earns its cost. Train it on carefully selected data from the brain, test it hard, serve it under your control, and feed the outcomes back so the brain and the model both improve.
This guide walks the ten steps in order. At each one it says plainly what Swfte provides today and what it is designed for. The short version: the identity graph, model hosting, the gateway and the governance vocabulary are built; document search and the sanitisation gateway are in progress; managed training, data export from the brain and domain evaluation suites are on the roadmap. The company brain overview and custom models overview give the full product picture.
Why start with the brain rather than the model?
Because a model trained on data you cannot trace is a model you cannot defend. A company brain holds what the organisation knows, who owns what and how things are connected, with evidence behind each fact and history you can read as of a given date. When training data later comes out of that brain, every example can carry its source, its access rules and the purpose it was selected for. Starting with the model inverts that, and teams usually discover the cost when an auditor asks where an example came from.
Step 1: Connect identity and structure
Everything downstream depends on knowing who people are, which groups they belong to and who reports to whom. That is what decides which content a person, or an agent acting for them, may see.
Built today. Swfte Enterprise Intelligence, the customer-hosted appliance behind the company brain, syncs from Active Directory or LDAP, Microsoft Entra ID, Okta and Google Workspace with a read-only account. It resolves accounts into one person with evidence on each link, models groups with transitive membership, reporting lines and org units, and keeps service accounts and devices as what they are. Each fact carries one of seven evidence statuses, from observed to unknown. A local API serves the graph, people, groups, coverage and audit, with the tenant always taken from the credential. Passwords and credentials are never read or stored.
Step 2: Add document and system sources
Identity tells you who. Documents, tickets, code and business systems tell you what the organisation actually knows and does.
In progress. The content store that keeps a per-object access list copied from the source, along with chunking and hybrid keyword and vector search, exists as libraries. Search filters by the asker's principals inside the database, so an object with an unknown access list stays invisible. The HTTP endpoints for search and context are not finished.
On the roadmap. Collectors for cloud, code, CI/CD, Kubernetes, databases, SaaS and business systems, tickets, and document and email connectors.
Step 3: Use the brain for retrieval and agents first
Before you think about training, put the brain to work. Most questions people ask of an organisation are knowledge questions, and retrieval answers those better than any tuned model, because the answer stays current when the source changes. Agents that look up owners, route requests or assemble context get most of their value from good retrieval and correct access rules.
The rule that matters here is that an agent never has more access than the person it acts for, and retrieved content is treated as untrusted data that cannot override policy. See agents on the brain for how that works.
Built today. The Trust Profile and the five autonomy levels, from L1 Assist to L5 Adaptive, give you the vocabulary for how far an agent may act on its own. On the roadmap. Wiring the brain into Cortex, Nexus and Studio, the context-package API and MCP server, and an action gateway with approvals. Today those products do not read from the brain.
Step 4: Find where behaviour, not knowledge, is the limit
Run your agents and retrieval for a while and watch where they fail. Most failures will be missing or stale content, which you fix in the brain. A few will be different: the right context reaches the model and it still gets the task wrong in the same way, again and again.
A worked example. A support organisation has its product documentation and past tickets in the brain. Retrieval brings the right article to the model every time. Yet the model's triage labels keep drifting from the organisation's own taxonomy, it writes replies in a register the team does not use, and it skips a mandatory escalation note about one time in a batch. None of that is a knowledge gap. It is a behaviour gap on a stable, narrow, high-volume task, and that is the shape a custom model is good at. Our post on when to fine-tune sets out the readiness tests in more detail.
Step 5: Select and sanitise training data, with consent and purpose recorded
Now you need examples of the task done right. The brain is a good source because it already knows where each item came from and who may see it.
Selection means choosing examples that cover the hard cases, resolving labelling disagreements in a written guideline, and keeping a held-out evaluation set apart from the start. Sanitisation means removing or minimising personal and restricted data before anything leaves its source. Recording consent and purpose means writing down, for each dataset, why it was assembled and under what basis, so the record exists when someone asks.
Built today. Personal and restricted data stay local by default in the appliance, secret-looking values are removed before storage, and per-person export and erase commands exist. In progress. The pre-model sanitisation gateway. On the roadmap. Data export from the brain for training and consent capture. The data preparation page covers the method.
Step 6: Adapt an open-weight base
Pick an open-weight base whose licence suits your use, reading the licence file shipped with the exact weights. Permissive licences such as Apache 2.0 and MIT are usually the simpler starting point. Start with a lightweight adapter, such as LoRA, before considering a full fine-tune, and keep a record of the base, the data version and the training settings.
On the roadmap. Managed adaptation runs on Swfte. The supported base models and methods for that service are not published yet. Today you train with your own tooling or a partner, then bring the weights to Swfte.
Step 7: Evaluate against domain and safety suites, with a regression gate
A tuned model is a new model. Score it against the held-out domain set and against the base or the current production model. Then run safety tests: refusal of harmful requests, over-refusal of legitimate ones, prompt injection through documents and tool output, and the same checks in every language you serve. Fine-tuning can weaken safety behaviour even on clean data, so the gate must block a candidate that regresses on safety, however good its task score.
Built today. Single-turn evaluation in Studio, and the published method for testing open-weight models. In progress. An agent evaluation sandbox with recorded, replayable traces, as internal tooling. On the roadmap. Named domain evaluation suites. More on the approach at evaluation and safety.
Step 8: Red-team it
Evaluation measures known risks. Red-teaming looks for the ones you did not think of: people trying to talk the model out of its rules, inputs crafted to exploit the tuned behaviour, misuse that only appears in the context of your workflows. Write down what was tried and what was found, and fix or disclose each finding before release.
On the roadmap. An automated red-team pipeline. Today this is a manual exercise you run with your own team or a specialist.
Step 9: Deploy behind Connect and bring agents onto it
Upload the approved weights into Model Vault. It records a checksum manifest and the base model, keeps versions, promotes them through development, staging and production, and deploys to a dedicated endpoint with an audit log you can export. Put Connect in front, so agents and applications reach the custom model through the same OpenAI-compatible API they already use, with routing, failover, usage caps, cost tracking and audit events. Route the narrow task to your model and leave everything else on the models that handle it well.
Built today. Model Vault and Connect, including content-policy checks with redaction at the gateway. Bring agents onto the new model one at a time, at an autonomy level that matches the evidence you have.
Step 10: Write outcomes back and decide when to retrain
The final step closes the loop. Every decision an agent makes with the custom model has an outcome: a ticket resolved or reopened, a summary approved or corrected. Those outcomes are evidence. Written back into the brain, they show where the model is drifting, which cases it struggles with and which corrections people keep making. That evidence becomes the next round of training examples and the trigger for retraining.
On the roadmap. Write-back into the brain, drift detection and retraining triggers. Today you capture outcomes in your own systems and decide on retraining by review.
What does the closed loop look like?
The loop has four stages: the company brain holds what the organisation knows; a custom model learns the behaviour your narrow task needs; agents use both to do the work under policy; outcomes flow back into the brain as new evidence. Each pass leaves the brain better informed and the model better targeted. The traceability runs the whole way, from data to model to agent to decision to action to outcome, which is what lets you answer for any single result.
| Stage | Built today | In progress | Roadmap |
|---|---|---|---|
| Company brain | Directory graph, local API, audit | Document search, sanitisation gateway | Other collectors, data export |
| Custom model | Model Vault, Connect | Agent eval sandbox, internal | Managed adaptation, domain eval suites |
| Agents | Trust Profile, autonomy levels | Brain wiring into Cortex, Nexus, Studio | |
| Outcomes | Audit logs | Write-back, retraining triggers |
The idea of a learning organisation built on this loop is explored further in the intelligence loop and in building AI agents from your own data.
Where should you begin?
Begin at step one, and do not skip ahead. Connecting identity and structure is useful the day it is done, because it makes access rules correct for everything that follows. Retrieval and agents on the brain solve most problems without training anything. The custom model comes later, for the one or two tasks where behaviour is the real limit, and by then you will have the data, the evidence and the evaluation set to do it properly.