Platform / Custom models / Data preparation

Data preparation for custom models: selecting, sanitising and recording what a model learns from

How to choose training data by purpose and by access, remove what must never reach a model, keep evaluation honest, and keep a record you can defend later.

A custom model is mostly its data. The base model, the method and the hardware matter, but the examples decide what the model learns to do and what it learns to repeat. This page covers sourcing from the company brain, sanitisation before any model sees a record, consent and lawful basis, cleaning and splits that do not leak, provenance, and what legal holds and erasure mean for a trained model. It is honest about what the brain exports today, which for training is nothing yet.

Why the data decides the result

A model trained on a few thousand carefully chosen, correctly labelled examples usually beats one trained on a large pile of whatever was easy to export. The model cannot tell a reviewed answer from a hurried one, a current policy from a withdrawn one, or a fact someone confirmed from a guess someone typed. It learns all of them with equal confidence.

So most of the work in a custom model is choosing what goes in and being able to say why. That work also produces the evidence a reviewer, an auditor or your own team will ask for later: which records were used, under what basis, with what removed, and how the evaluation set was kept apart. If you cannot answer those questions, the model cannot be defended, however well it scores.

Sourcing from the company brain

The brain is designed to be the place training data is chosen from, because it already knows what each fact rests on and who may see it.

  • Select by purpose

    Start from the task in the brief and pull only the records that serve it. A model that classifies correspondence does not need the reporting lines of the whole organisation.

  • Select by access

    Training data should respect the same rule as every other read: never more access than the people the model will serve. Records with incomplete access lists stay out.

  • Labels from strong evidence

    Every fact in the brain carries one of seven evidence statuses. Use verified or corroborated facts as labels. Inferred, stale, disputed and unknown facts are not ground truth.

  • History for as-of correctness

    The brain keeps insert-only history, so an example can be paired with the organisation as it stood when the decision was made, not as it stands today.

Sanitisation before any model sees a record

Sanitisation happens before training, not after. The pre-model sanitisation gateway, which is in progress, sits between stored content and any model and removes what must not reach one. The brain already refuses to read or store passwords and other credentials, and removes values that look like secrets before anything is stored. Personal and restricted data are local-only by default.

For training data the gateway is designed to go further: strip or replace direct identifiers, mask account numbers and similar values, and drop records that are flagged restricted for this purpose. Each removal is recorded, so the datasheet can say what was taken out and by which rule. Until the gateway is finished, teams preparing data with their own pipeline should apply the same idea: sanitise first, and keep the log of what was removed.

Cleaning, deduplication, decontamination and splits

The mechanical steps, in order. Skipping any of them makes the evaluation result unreliable.

  1. 01

    Clean

    Remove empty, truncated and malformed records, normalise encodings and formats, and fix labels that reviewers have since corrected.

  2. 02

    Deduplicate

    Exact and near-duplicate examples teach the model to repeat themselves and inflate scores. Keep one of each.

  3. 03

    Split by unit, not by row

    Split by case, customer or thread, so two halves of one conversation never land on opposite sides. That is how leaks happen.

  4. 04

    Decontaminate

    Check every training example against the evaluation set and remove anything that matches or nearly matches. A model that saw the test is not being tested.

  5. 05

    Treat synthetic data with care

    Generated examples can fill gaps, but they carry the generating model's habits and errors, and its terms may restrict this use. Label them as synthetic and never evaluate on them alone.

Provenance and a datasheet

Every dataset should leave with a datasheet: a short document that says where the records came from, the date range, the selection rule, what sanitisation removed, the evidence statuses accepted as labels, the lawful basis your team recorded, the split method and the decontamination check. It also names who assembled it and who approved it. A hash of the final files ties the datasheet to exactly what was trained on.

The datasheet belongs in the model's lineage record, next to the base model and version the Model Vault records and the evaluation results that justified promotion. When someone asks months later why the model behaves the way it does, the answer starts here.

Legal holds, erasure and what they mean for retraining

The brain supports per-person export and erase commands and optional retention that pseudonymises deleted people; legal holds on content are part of its knowledge core, which is in progress. A held record must not be deleted; an erased record must not come back. Both affect training. A dataset that contains a person later erased is no longer a dataset you can rebuild as it was.

For a trained model the honest position is that removing one person's influence from trained weights is not something current methods do reliably. The practical answer is to rebuild the dataset without the erased records and retrain, then evaluate and promote the new version. Plan for this from the start: keep the pipeline repeatable, keep the datasheet, and decide in advance how quickly a retrain follows an erasure request.

What you can do today, and a worked example

The brain does not export training data today; export for training is on the roadmap. What you can do today is prepare data with your own pipeline, following the steps on this page, train with your own tooling, and host the result in the Model Vault, where the base model and versions are recorded and the endpoint sits behind Connect. Supported adaptation methods for a managed service are <supported adaptation methods - founder to fill>.

A worked example. An internal service desk wants a model that drafts first replies in its own format. The team takes resolved tickets whose resolutions a lead had confirmed, drops anything with restricted fields, masks names and account numbers, removes near-duplicates and splits by ticket thread. It sets aside the most recent month as the evaluation set and checks no training example matches it. The datasheet records every step and the basis the legal team approved. When a former employee later asks to be erased, the team rebuilds the dataset without those tickets and retrains.

Data preparation: built, in progress and roadmap

What is built, in progress and on the roadmap: Data preparation
CapabilityStatusNotes
Seven evidence statuses on brain factsBuiltObserved, corroborated, verified, inferred, stale, disputed, unknown.
Insert-only history and as-of readsBuiltPair examples with the organisation as it stood at the time.
Credentials never read or storedBuiltSecret-looking values are removed before storage.
Per-person export and eraseBuiltWith optional retention that pseudonymises deleted people.
Per-object access lists on contentIn progressExists as libraries; the search endpoints are not finished.
Pre-model sanitisation gatewayIn progressNot yet something to rely on.
Training data export from the brainRoadmapNothing is exported for training today.
Consent and purpose captureRoadmapDesigned to record the decision your legal team makes.

Legend

  • Built. Exists today and can be used.
  • In progress. Being built. Not yet something to rely on.
  • Roadmap. Designed for and on the roadmap. Not built. No dates are given.

Where this fits in the loop

Data preparation is the edge from the company brain to the custom model: select, sanitise, then adapt.

The same four stations and four arrows are listed in order below.
  1. 01Company brainHolds what the organisation knows, with evidence statuses, history and access rules.(this page)
  2. 02Custom modelAdapted on data chosen from the brain, then evaluated and hardened before it ships.(this page)
  3. 03Governed agentsUse the model and read the brain, inside a Trust Profile, with approval where it matters.
  4. 04OutcomesWhat happened: approvals, corrections, results and cost, all on the record.

The four arrows

  1. Company brain to Custom model: select, sanitise, adaptRoadmap

    Choose training data from the brain, remove what must not reach a model, adapt an open-weight base. The sanitisation gateway is in progress, and the data selection and training steps are on the roadmap.

  2. Custom model to Governed agents: serve, governBuilt

    Serve the model on dedicated infrastructure behind the Connect gateway and bring agents onto it under policy. Model hosting and the gateway are built.

  3. Governed agents to Outcomes: act, recordBuilt

    Agents act within their Trust Profile, with human approval for consequential steps, and every action is recorded.

  4. Outcomes to Company brain: written back as evidenceRoadmap

    Outcomes return to the brain as new evidence with a status, and they decide when the model needs retraining. The write-back is on the roadmap.

Legend

  • Built. Exists today and can be used.
  • In progress. Being built. Not yet something to rely on.
  • Roadmap. Designed for and on the roadmap. Not built. No dates are given.

Frequently asked questions

Can we export training data from the company brain today?

No. Export for training is on the roadmap. Today you prepare data with your own pipeline, train with your own tooling, and host the resulting weights in the Model Vault, where the base model and versions are recorded and Connect serves the endpoint.

Which facts should be used as training labels?

Facts with a strong evidence status, such as verified or corroborated. Inferred, stale, disputed and unknown facts are not ground truth, and a model trained on them learns their uncertainty as if it were certain.

What does sanitisation remove?

The brain already never stores passwords or credentials and removes secret-looking values. The pre-model sanitisation gateway, in progress, is designed to strip direct identifiers, mask values such as account numbers and drop records restricted for the purpose, logging every removal for the datasheet.

Is training on our data GDPR compliant?

Swfte does not claim any training use is compliant, and does not give legal advice. Whether records may be used is your organisation's decision with its legal team. The platform is designed to record purpose and basis as evidence that helps you meet your own obligations; the posture depends on use case, jurisdiction, deployment and configuration.

What is decontamination?

Checking every training example against the evaluation set and removing anything that matches or nearly matches. If the model trained on the test cases, its score tells you about memory, not about how it will handle new work.

What happens to a model if a person asks to be erased?

Current methods cannot reliably remove one person from trained weights. The honest answer is to rebuild the dataset without their records, retrain, evaluate and promote the new version. Keep the pipeline repeatable so this is routine rather than a crisis.

Is synthetic data a shortcut?

Sometimes useful, never a shortcut. Generated examples carry the generating model's habits and errors, and its terms may restrict the use. Label them as synthetic, mix them with real examples, and never evaluate on synthetic data alone.

Take data preparation further with Swfte

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.