Platform / Custom models / Data preparation
Data preparation for custom models: selecting, sanitising and recording what a model learns from
How to choose training data by purpose and by access, remove what must never reach a model, keep evaluation honest, and keep a record you can defend later.
A custom model is mostly its data. The base model, the method and the hardware matter, but the examples decide what the model learns to do and what it learns to repeat. This page covers sourcing from the company brain, sanitisation before any model sees a record, consent and lawful basis, cleaning and splits that do not leak, provenance, and what legal holds and erasure mean for a trained model. It is honest about what the brain exports today, which for training is nothing yet.
Why the data decides the result
A model trained on a few thousand carefully chosen, correctly labelled examples usually beats one trained on a large pile of whatever was easy to export. The model cannot tell a reviewed answer from a hurried one, a current policy from a withdrawn one, or a fact someone confirmed from a guess someone typed. It learns all of them with equal confidence.
So most of the work in a custom model is choosing what goes in and being able to say why. That work also produces the evidence a reviewer, an auditor or your own team will ask for later: which records were used, under what basis, with what removed, and how the evaluation set was kept apart. If you cannot answer those questions, the model cannot be defended, however well it scores.
Sourcing from the company brain
The brain is designed to be the place training data is chosen from, because it already knows what each fact rests on and who may see it.
Select by purpose
Start from the task in the brief and pull only the records that serve it. A model that classifies correspondence does not need the reporting lines of the whole organisation.
Select by access
Training data should respect the same rule as every other read: never more access than the people the model will serve. Records with incomplete access lists stay out.
Labels from strong evidence
Every fact in the brain carries one of seven evidence statuses. Use verified or corroborated facts as labels. Inferred, stale, disputed and unknown facts are not ground truth.
History for as-of correctness
The brain keeps insert-only history, so an example can be paired with the organisation as it stood when the decision was made, not as it stands today.
Sanitisation before any model sees a record
Sanitisation happens before training, not after. The pre-model sanitisation gateway, which is in progress, sits between stored content and any model and removes what must not reach one. The brain already refuses to read or store passwords and other credentials, and removes values that look like secrets before anything is stored. Personal and restricted data are local-only by default.
For training data the gateway is designed to go further: strip or replace direct identifiers, mask account numbers and similar values, and drop records that are flagged restricted for this purpose. Each removal is recorded, so the datasheet can say what was taken out and by which rule. Until the gateway is finished, teams preparing data with their own pipeline should apply the same idea: sanitise first, and keep the log of what was removed.
Consent and lawful basis: your decision, recorded
Whether a given set of records may be used to train a model is a decision for your organisation and its legal team. It depends on what the people concerned were told, what basis the data was collected under, the jurisdiction and the purpose. Swfte does not give legal advice and does not decide this for you.
What the platform is designed to do is record the decision: the purpose a dataset was assembled for, the lawful basis your team recorded, and any consent conditions that apply, so the record travels with the dataset and the model trained on it. Consent capture is on the roadmap. The goal is technical controls and evidence that help an organisation meet its own obligations, with the posture depending on use case, jurisdiction, deployment and configuration.
Cleaning, deduplication, decontamination and splits
The mechanical steps, in order. Skipping any of them makes the evaluation result unreliable.
- 01
Clean
Remove empty, truncated and malformed records, normalise encodings and formats, and fix labels that reviewers have since corrected.
- 02
Deduplicate
Exact and near-duplicate examples teach the model to repeat themselves and inflate scores. Keep one of each.
- 03
Split by unit, not by row
Split by case, customer or thread, so two halves of one conversation never land on opposite sides. That is how leaks happen.
- 04
Decontaminate
Check every training example against the evaluation set and remove anything that matches or nearly matches. A model that saw the test is not being tested.
- 05
Treat synthetic data with care
Generated examples can fill gaps, but they carry the generating model's habits and errors, and its terms may restrict this use. Label them as synthetic and never evaluate on them alone.
Provenance and a datasheet
Every dataset should leave with a datasheet: a short document that says where the records came from, the date range, the selection rule, what sanitisation removed, the evidence statuses accepted as labels, the lawful basis your team recorded, the split method and the decontamination check. It also names who assembled it and who approved it. A hash of the final files ties the datasheet to exactly what was trained on.
The datasheet belongs in the model's lineage record, next to the base model and version the Model Vault records and the evaluation results that justified promotion. When someone asks months later why the model behaves the way it does, the answer starts here.
Legal holds, erasure and what they mean for retraining
The brain supports per-person export and erase commands and optional retention that pseudonymises deleted people; legal holds on content are part of its knowledge core, which is in progress. A held record must not be deleted; an erased record must not come back. Both affect training. A dataset that contains a person later erased is no longer a dataset you can rebuild as it was.
For a trained model the honest position is that removing one person's influence from trained weights is not something current methods do reliably. The practical answer is to rebuild the dataset without the erased records and retrain, then evaluate and promote the new version. Plan for this from the start: keep the pipeline repeatable, keep the datasheet, and decide in advance how quickly a retrain follows an erasure request.
What you can do today, and a worked example
The brain does not export training data today; export for training is on the roadmap. What you can do today is prepare data with your own pipeline, following the steps on this page, train with your own tooling, and host the result in the Model Vault, where the base model and versions are recorded and the endpoint sits behind Connect. Supported adaptation methods for a managed service are <supported adaptation methods - founder to fill>.
A worked example. An internal service desk wants a model that drafts first replies in its own format. The team takes resolved tickets whose resolutions a lead had confirmed, drops anything with restricted fields, masks names and account numbers, removes near-duplicates and splits by ticket thread. It sets aside the most recent month as the evaluation set and checks no training example matches it. The datasheet records every step and the basis the legal team approved. When a former employee later asks to be erased, the team rebuilds the dataset without those tickets and retrains.
Data preparation: built, in progress and roadmap
| Capability | Status | Notes |
|---|---|---|
| Seven evidence statuses on brain facts | Built | Observed, corroborated, verified, inferred, stale, disputed, unknown. |
| Insert-only history and as-of reads | Built | Pair examples with the organisation as it stood at the time. |
| Credentials never read or stored | Built | Secret-looking values are removed before storage. |
| Per-person export and erase | Built | With optional retention that pseudonymises deleted people. |
| Per-object access lists on content | In progress | Exists as libraries; the search endpoints are not finished. |
| Pre-model sanitisation gateway | In progress | Not yet something to rely on. |
| Training data export from the brain | Roadmap | Nothing is exported for training today. |
| Consent and purpose capture | Roadmap | Designed to record the decision your legal team makes. |
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Where this fits in the loop
Data preparation is the edge from the company brain to the custom model: select, sanitise, then adapt.
- 01Company brainHolds what the organisation knows, with evidence statuses, history and access rules.(this page)
- 02Custom modelAdapted on data chosen from the brain, then evaluated and hardened before it ships.(this page)
- 03Governed agentsUse the model and read the brain, inside a Trust Profile, with approval where it matters.
- 04OutcomesWhat happened: approvals, corrections, results and cost, all on the record.
The four arrows
- Company brain to Custom model: select, sanitise, adaptRoadmap
Choose training data from the brain, remove what must not reach a model, adapt an open-weight base. The sanitisation gateway is in progress, and the data selection and training steps are on the roadmap.
- Custom model to Governed agents: serve, governBuilt
Serve the model on dedicated infrastructure behind the Connect gateway and bring agents onto it under policy. Model hosting and the gateway are built.
- Governed agents to Outcomes: act, recordBuilt
Agents act within their Trust Profile, with human approval for consequential steps, and every action is recorded.
- Outcomes to Company brain: written back as evidenceRoadmap
Outcomes return to the brain as new evidence with a status, and they decide when the model needs retraining. The write-back is on the roadmap.
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Frequently asked questions
Can we export training data from the company brain today?
No. Export for training is on the roadmap. Today you prepare data with your own pipeline, train with your own tooling, and host the resulting weights in the Model Vault, where the base model and versions are recorded and Connect serves the endpoint.
Which facts should be used as training labels?
Facts with a strong evidence status, such as verified or corroborated. Inferred, stale, disputed and unknown facts are not ground truth, and a model trained on them learns their uncertainty as if it were certain.
What does sanitisation remove?
The brain already never stores passwords or credentials and removes secret-looking values. The pre-model sanitisation gateway, in progress, is designed to strip direct identifiers, mask values such as account numbers and drop records restricted for the purpose, logging every removal for the datasheet.
Is training on our data GDPR compliant?
Swfte does not claim any training use is compliant, and does not give legal advice. Whether records may be used is your organisation's decision with its legal team. The platform is designed to record purpose and basis as evidence that helps you meet your own obligations; the posture depends on use case, jurisdiction, deployment and configuration.
What is decontamination?
Checking every training example against the evaluation set and removing anything that matches or nearly matches. If the model trained on the test cases, its score tells you about memory, not about how it will handle new work.
What happens to a model if a person asks to be erased?
Current methods cannot reliably remove one person from trained weights. The honest answer is to rebuild the dataset without their records, retrain, evaluate and promote the new version. Keep the pipeline repeatable so this is routine rather than a crisis.
Is synthetic data a shortcut?
Sometimes useful, never a shortcut. Generated examples carry the generating model's habits and errors, and its terms may restrict the use. Label them as synthetic, mix them with real examples, and never evaluate on synthetic data alone.
Take data preparation further with Swfte
Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.