Platform / Custom models / Deploy and serve
Deploy and serve a custom model on infrastructure you control
How a custom model gets from a weights file to a governed endpoint: the Model Vault, the Connect gateway, promotion, rollback and the policy around every request.
This is the part of custom models that Swfte has built. You bring weights you have adapted, the Model Vault records and versions them, a dedicated endpoint serves them, and Connect puts one governed API in front, with routing, usage caps, cost tracking and audit events. This page walks through where the model runs, the vault flow step by step, serving engines and quantisation, rollback, gateway policy, monitoring, and EU-first hosting, and it says which parts are still on the roadmap.
Where the model runs
The choice of where depends on who must be able to see prompts, weights and logs. The deploy-models guides go deeper on each option.
Model Vault, dedicated endpoint
Your weights deployed to a dedicated endpoint, not a shared pool. Built today, with an OpenAI-compatible API mode and a health endpoint.
Read more about Model Vault, dedicated endpointConnect in front
Connect fronts any OpenAI-compatible endpoint, including one you already run, so the governed path does not depend on where the model sits.
Read more about Connect in frontDedicated cloud
Single-tenant infrastructure in a region you choose, for teams that want Swfte to run the stack without sharing it.
Read more about Dedicated cloudSelf-deployed
Run the gateway and models in your own environment. Start from the self-deploy page and the self-hosted inference guide.
Read more about Self-deployedOn-premises and air-gapped
Designed for, and available on request. An air-gapped install for model serving is not shipped as a standard product today.
Read more about On-premises and air-gapped
The vault flow in plain words
Every step leaves a record. Together they are how you answer the question: what exactly is serving this request?
- 01
Upload the weights
Upload your weights as a multipart upload, so files of many gigabytes are handled as a matter of course.
- 02
Hash manifest
A sha256 hash is recorded for every file, so the weights you deploy can be checked against the weights you evaluated.
- 03
Record the base model
The vault records the parent model a fine-tune was built from. The base licence still applies, and this is where you can show which one.
- 04
Version
Each upload is a version within a group, so a new training run is a new version, not an overwrite.
- 05
Promote
Move the version from development to staging to production, and archive it when it retires. Promotion should follow the evaluation gate.
- 06
Deploy
Deploy the promoted version to a dedicated endpoint and undeploy it when it is no longer needed.
- 07
Audit
Every action on the model is written to its audit log, which can be exported for review.
Serving engines and quantisation
The serving engine decides how many requests one GPU can carry. For shared serving, engines such as vLLM and SGLang keep the card busy with continuous batching; Ollama and llama.cpp suit development and single-team tools. The trade-offs are set out on the deploy-models hub, and the choice is made per deployment against the model, the hardware and the traffic.
Quantisation shrinks the weights so a model fits on less memory, and it changes the model. The evaluation you ran on full-precision weights says nothing certain about the 8-bit or 4-bit file you actually serve, especially on refusal behaviour and on languages other than English. Treat each quantised artifact as a new candidate: re-run the evaluation on exactly the file you will deploy, and record its hash. If you serve an adapter on top of a base, test that combination. How Swfte serves adapters is <how adapters are served - founder to fill>.
Routing and rollback through Connect
Connect is one OpenAI-compatible API in front of many providers. A custom endpoint sits behind it like any other model, so agents and applications call one address and the gateway decides where each request goes. That makes routing policy, not code: send the high-volume, low-risk work to your custom model, keep a fallback for when it is unavailable, and change the split by configuration.
Routing is also the rollback lever. Keep the previous approved revision deployed and routable until the new one has run on live traffic for long enough to trust. If something goes wrong, rolling back is a routing change and a log entry, not an emergency redeploy. Once you are confident, undeploy the old revision and archive it in the vault.
Policy at the gateway
The gateway is the one place every request passes through, so it is where policy is enforced.
Approved-model policy
Routing rules decide which models production traffic can reach. Point production routes only at versions you have promoted, so an unapproved model is not one setting away from serving people.
Usage caps
Workspace and per-model caps on tokens and spend, so a runaway loop cannot consume the budget for everyone.
Cost tracking
Usage and spend recorded through the gateway, so you can see what the custom model costs next to the alternatives.
Content policy
Secret and personal-data detectors, with a redact action, applied before content reaches the model.
Audit events
A stream of events for every request class, so you can show which model served what, and under which policy.
Monitoring and retraining triggers
Today usage, cost and latency are visible through the gateway's usage and analytics reads, and the audit stream records what was served. That is enough to see when a model is slow, expensive or suddenly busier than planned, and usage caps can act on it before the budget is gone. It is not the same as knowing whether the answers are still good.
What is not built yet is automatic detection of quality drift: noticing that a model's answers have become less accurate as the work, the vocabulary or the underlying facts change, and raising a retraining trigger. That is designed for and on the roadmap. Until then, run your domain suite on a schedule against the production model, sample live outputs for review, and treat a falling score as a reason to plan the next training run.
EU-first hosting
For European organisations the default should be that weights, prompts, outputs and logs stay in the EU, on infrastructure where you know who operates it and under which law. Dedicated, in-region hosting is the stated direction of the platform; the regions and facilities on offer are <EU regions and facilities offered - founder to fill>. The guide to deploying an LLM in the EU explains the difference between residency and sovereignty, and why a region label alone is not enough.
Hosting in region is a technical control, and a record of where each model version ran is evidence. Together they help an organisation meet its own obligations; they do not settle them. The posture depends on use case, jurisdiction, deployment and configuration, so decide the hosting option with the people who own those obligations, and keep the decision with the model's lineage record.
A worked example, and the honest status
A legal operations team has trained an adapter on its own infrastructure and merged it into a small open-weight base. It uploads the merged weights to the Model Vault, which records the hashes and the base model, and creates version one in development. It quantises to 8-bit for serving, uploads that as a separate version and runs the evaluation suite on the quantised file. The results hold, so the version is promoted to staging, then to production after sign-off, and deployed to a dedicated endpoint.
In Connect, the team routes contract-summary requests to the new endpoint and keeps the general model it used before as the fallback. Usage caps are set for the workspace and the model. Some weeks later a reviewer notices weaker answers on a new contract type; the team routes that request class back to the fallback while it prepares a retrain. Hosting, versioning, promotion, deployment, gateway routing and audit are built. Managed training, drift detection and adapter serving are not.
Deploy and serve: built, in progress and roadmap
| Capability | Status | Notes |
|---|---|---|
| Model Vault upload with sha256 manifest | Built | Multipart upload, versions and base model recorded. |
| Stage promotion and audit log | Built | Development, staging, production, archived. Audit log can be exported. |
| Deploy to a dedicated endpoint | Built | OpenAI-compatible API mode and a health endpoint. |
| Connect gateway in front of custom endpoints | Built | Routing, failover, usage caps, cost tracking and audit events. |
| Content policy with redaction | Built | Secret and personal-data detectors at the gateway. |
| Air-gapped model serving install | Roadmap | Designed for; on request. Not shipped as a standard install. |
| Adapter serving as adapters | Roadmap | Serve merged weights today. |
| Quality drift detection and retraining triggers | Roadmap | Usage, cost and latency are visible today. |
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Where this fits in the loop
Deploy and serve is the edge from custom model to governed agents: the model is served behind Connect, and agents reach it only under policy.
- 01Company brainHolds what the organisation knows, with evidence statuses, history and access rules.
- 02Custom modelAdapted on data chosen from the brain, then evaluated and hardened before it ships.(this page)
- 03Governed agentsUse the model and read the brain, inside a Trust Profile, with approval where it matters.(this page)
- 04OutcomesWhat happened: approvals, corrections, results and cost, all on the record.
The four arrows
- Company brain to Custom model: select, sanitise, adaptRoadmap
Choose training data from the brain, remove what must not reach a model, adapt an open-weight base. The sanitisation gateway is in progress, and the data selection and training steps are on the roadmap.
- Custom model to Governed agents: serve, governBuilt
Serve the model on dedicated infrastructure behind the Connect gateway and bring agents onto it under policy. Model hosting and the gateway are built.
- Governed agents to Outcomes: act, recordBuilt
Agents act within their Trust Profile, with human approval for consequential steps, and every action is recorded.
- Outcomes to Company brain: written back as evidenceRoadmap
Outcomes return to the brain as new evidence with a status, and they decide when the model needs retraining. The write-back is on the roadmap.
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Frequently asked questions
Can we host our own fine-tuned weights on Swfte today?
Yes. The Model Vault takes your weights, records a sha256 manifest and the base model, versions them, promotes them through development, staging and production, and deploys them to a dedicated endpoint. Connect then serves them through one OpenAI-compatible API.
Do we have to rewrite our applications to use a custom model?
No. Connect exposes one OpenAI-compatible API in front of every model, including your custom endpoint. Applications call the gateway, and routing decides which model answers, so moving traffic is a configuration change.
How do we roll back a bad release?
Keep the previous approved revision deployed and routable until the new one has proven itself on live traffic. Rolling back is then a routing change in Connect plus an audit entry, rather than an urgent redeploy.
Why re-evaluate a quantised model?
Because quantisation changes the model. Results on full-precision weights do not reliably carry over to an 8-bit or 4-bit file, especially for refusal behaviour and non-English text. Test the exact artifact you serve, and record its hash.
Does Swfte detect when our model starts getting worse?
Usage, cost and latency are visible through the gateway today. Automatic quality drift detection and retraining triggers are on the roadmap. Until then, run your domain suite against production on a schedule and review samples of live output.
Is hosting a model in the EU on Swfte GDPR compliant?
Swfte does not claim any deployment is compliant. In-region, dedicated hosting and gateway controls are technical measures and evidence that help an organisation meet its own obligations. The posture depends on use case, jurisdiction, deployment and configuration.
Can we run the model on our own premises or air-gapped?
On-premises and air-gapped serving are designed for and available on request. An air-gapped install for model serving is not shipped as a standard product today. Self-deployment of the gateway and dedicated cloud options are described on their own pages.
Take deploy and serve further with Swfte
Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.