← The journal
Guides

Building a Company Brain Without Losing Control of Your Data

How to build a company brain and keep control: where it runs, read-only access, identity-aware reads, audit.

Swfte Journal / Guides

You can build a company brain without losing control of your data if you settle a few decisions before connecting anything: where it runs, who owns each source, and which classes of data may never leave. Then start with identity, connect every source read-only, never store credentials, filter every read by the person asking, sanitise content before it reaches a model, keep a kill switch within reach, and make the audit trail something you can verify yourself. Add sources one family at a time.

A company brain is one governed place holding an organisation's sources of information as an evidence-backed, time-aware graph that AI products read from. If that definition is new, start with what is a company brain for AI. This guide assumes you want one and are worried, reasonably, about putting the organisation's most sensitive context in one place.

What should you decide before connecting anything?

Three decisions shape everything after them. Write the answers down and get them signed off by whoever owns data protection and security.

Where it runs. A brain can run in one of three modes. Connected: it runs in your environment and keeps an outbound-only link to a vendor control plane for updates and management. Private: you host the control plane too. Air-gapped: no outbound connection at all, with updates carried in by hand. Pick the mode by the most sensitive source you expect to connect, not the first one. The deployment page describes each mode and what it costs you in convenience.

Who owns each source. For each system you plan to connect, name the person who can approve the connection and who will be asked when something looks wrong. If nobody will put their name to a source, do not connect it yet.

What may never leave. List the classes of data that must stay inside your environment whatever happens: personal data about staff, restricted data such as health or legal matters, anything under a contractual residency clause. These become rules the brain enforces, not guidance people are expected to remember. Our post on data sovereignty for AI is a good primer on why location alone does not settle this.

Why start with identity and the directory?

Every later question needs to know who someone is. Who owns this, who may read this, who should approve this: each is a question about people and groups. If the brain cannot tell that three accounts belong to one person, it will get every one of those questions wrong.

The directory is also the safest place to start. It holds structure (accounts, groups, reporting lines, org units) rather than the contents of documents or messages. It lets you test the whole operating model, from installation to audit to erase, on data that is sensitive but bounded.

In Swfte's case, this is the part that is built. The appliance syncs from Active Directory and LDAP, Microsoft Entra ID, Okta and Google Workspace, resolves accounts into people with the evidence for each link, works out transitive group membership, and keeps service accounts and devices as what they are rather than counting them as staff. The sources page lists each source and its status.

How should the brain connect to sources?

Three rules, and none of them should be negotiable.

Read-only accounts. The connector's account should be able to read and nothing else. If a vendor's setup guide asks for write or admin rights to sync a directory, ask why.

Attribute allow-lists. Do not sync every attribute a directory holds. Decide which fields the brain needs, such as name, email, department, manager and group memberships, and allow only those. Home addresses, phone numbers and custom HR fields can stay where they are.

Never credentials. Passwords, password hashes and secrets have no place in a company brain. In Swfte's appliance the directory connector refuses to start if it is configured to read credentials, and values that look like secrets are removed before anything is stored.

Where should personal data live?

Personal data and restricted data should be local-only by default: held in the brain inside your environment and not sent anywhere else, including to the vendor's control plane. That is the default in Swfte's appliance. Anything that is to leave should leave because someone decided it could, for a stated purpose, and that decision should be recorded.

This matters most in connected mode. The link to the control plane exists for updates and management commands. It should carry health and version information, not the contents of your graph.

How do you make sure nobody sees more than they should?

The rule is simple to state: a reader never gets more access than the person asking. An assistant answering a question for a junior analyst sees what that analyst can see in the source system, and no more. An agent acting on someone's behalf carries a short-lived, scoped token that says whose behalf, and its reads are filtered accordingly.

Two details make the difference between a rule and a slogan. First, the filtering should happen as close to the data as possible, ideally inside the database query, so an application bug cannot widen it. Second, an object whose access list is unknown or incomplete should be invisible, not public. Swfte's search and delegation libraries are built this way and tested, and the endpoints that expose them are in progress. The access and governance page covers the design.

One more rule belongs here: content retrieved from the brain is untrusted input to a model. A document that says "ignore your instructions" must not be able to change policy.

What should happen before anything reaches a model?

Sanitisation. Between the brain and any model, a gateway should remove or mask what the model does not need: identifiers, secrets that slipped past earlier checks, fields marked restricted. What is removed should depend on where the model runs. A model in your own environment may be allowed more than one behind a third-party API.

Swfte's pre-model sanitisation gateway is in progress. Today, content-policy checks with secret and personal-data detectors and a redact action run at the Connect gateway for model traffic that passes through it.

How do you keep the power to stop it and check it?

A kill switch you control

Every connected appliance should have a local kill switch: a control in your environment, under your operators, that cuts the link to the vendor immediately. It should not depend on the vendor's cooperation, and using it should be journalled so you can show later when it was pulled and by whom.

Swfte's edge link is outbound-only over mutual TLS, accepts only signed, typed commands from an allow-list with no shell capability, and has a local, journalled kill switch. Updates are signed. These are built.

Audit you can verify

An audit log you have to take on trust is a log of the vendor's opinion. A verifiable one lets you check for yourself that nothing was removed or altered. The usual technique is hash-chaining: each entry includes a hash of the one before, with periodic signed checkpoints, so tampering breaks the chain in a way anyone can detect.

Swfte's graph store keeps a hash-chained audit log with signed checkpoints, and the local API serves it. History in the graph is insert-only, so you can read what the brain held on any past date. Ask any vendor whether you can verify their log without their help.

How do retention, export and erase work?

People have rights over data about them, and organisations have their own retention rules. The brain should support both without a support ticket.

  • Export. A command that produces everything the brain holds about one person.
  • Erase. A command that removes a person and records that it was done.
  • Retention. An optional policy that pseudonymises people after they are deleted at the source, so history stays usable without naming them.

Swfte's appliance has per-person export and erase commands and optional retention that pseudonymises deleted people. These are technical controls that help an organisation meet its own obligations; the posture depends on use case, jurisdiction, deployment and configuration.

What is a sensible staged plan?

  1. Decide. Mode, source owners, data that may never leave. Get sign-off.
  2. Install in your environment. Container image, installer, Helm chart or the air-gapped route. Check the install can be verified and uninstalled.
  3. Connect the directory, read-only, with an attribute allow-list. Review identity resolution and coverage with the people who know the organisation.
  4. Exercise the controls. Run an export and an erase on a test person. Pull the kill switch. Verify the audit chain.
  5. Add documents. With their access lists, filtered by the asker. In Swfte this is in progress.
  6. Add systems as collectors arrive. Cloud, code, CI/CD, Kubernetes, databases, SaaS and tickets are on Swfte's roadmap, one family at a time, each with a named owner.
  7. Point AI tools at the brain. Wiring into Swfte's assistants, agents and studio is designed for and on the roadmap.

For the wider stack this sits in, the sovereign AI reference architecture and the sovereignty page are useful companions.

What are the common mistakes?

  • Connecting everything at once. You cannot review identity resolution and twenty sources in the same week.
  • One powerful service account. Convenient on day one, and the reason the index ignores permissions forever after.
  • Treating "unknown access" as "everyone". It should mean nobody.
  • Choosing the mode by the first source. Choose by the most sensitive one you expect to add.
  • No named owner per source. Then nobody answers when two sources disagree.
  • Never testing erase. The first real request is the wrong time to discover it does not work.

What should you ask any vendor?

  • Where does the brain run, and what leaves my environment in each mode?
  • Which permissions does each connector need, and can I restrict attributes?
  • Are credentials ever read or stored?
  • Is every read filtered by the asking user, and where does that filtering happen?
  • What happens to an object with an unknown access list?
  • Can I verify the audit log myself?
  • Is there a local kill switch that works without you?
  • How do per-person export and erase work, and can I test them now?
  • Which parts are built today, and which are on your roadmap?

Swfte's answers, including what is not built yet, are on the company brain pages and the trust centre.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.