Sovereign inference

Deploy open-weight and private models on infrastructure you control

Pick, evaluate, harden, deploy and monitor open-weight models on dedicated, in-region infrastructure, with an evaluation gate in front of every promotion.

Open-weight models changed the sourcing question. You no longer have to send prompts to a model lab to get frontier-class capability: you can run the weights yourself, in a region you choose, behind policy you write. The hard part is not downloading a checkpoint. It is deciding which one is safe to run, proving it before it touches production, and being able to show what ran, when, and why. Swfte treats that as one loop: Pick, Evaluate, Harden, Deploy, Monitor.

Pick, Evaluate, Harden, Deploy, Monitor

One loop for every model, whether it is an open-weight family, a fine-tune of your own, or a model you bring.

  1. Pick

    Shortlist by task, language coverage, context length and licence, not by leaderboard rank. Read the licence file that ships with the exact checkpoint. Record the revision you chose.

  2. Evaluate

    Run your own tasks, a safety suite and an EU-language suite against the candidate. The result decides whether it is allowed to go further. See how we test open-source models.

  3. Harden

    Verify weights against pinned hashes, load safetensors only, switch off remote code execution, set runtime guardrails, and fix the system policy the model runs under.

  4. Deploy

    Serve on dedicated infrastructure in your region, behind the Connect gateway, with a quantization and parallelism plan sized to the hardware. Promotion is gated on the evaluation result.

  5. Monitor

    Watch quality, safety signals, latency and cost per completed task on live traffic. A regression is a reason to roll back to the previous approved revision, and every promotion and rollback is logged.

Open-weight model families you can deploy

Families are named here, not versions: checkpoints and licences change month to month. The licence column describes a pattern as of 2026-10-06. Always read the licence shipped with the exact file you deploy.

FamilyMakerLicence patternCheck before you deploy
LlamaMetaLlama Community Licence: a custom licence, not an OSI-approved open-source licence, with a 700 million monthly-active-user clause.The Llama 4 acceptable use policy states that rights for its multimodal models are not granted to individuals domiciled in, or companies headquartered in, the EU. Have counsel read the current licence and policy before any EU deployment.
QwenAlibabaApache 2.0 across its open-weight releases.Not every Qwen model is open: some flagship checkpoints are API-only. Confirm that downloadable weights exist for the one you want.
DeepSeekDeepSeekMIT on R1 and on recent releases; earlier releases such as V3 used DeepSeek’s own model licence.Licence differs by release. Self-hosting the weights keeps prompts off the lab’s hosted API, which is a different thing from using that API.
GemmaGoogleApache 2.0 from Gemma 4 onwards; earlier generations used Google’s own Gemma terms.Apache 2.0 grants no trademark rights, so you cannot ship a product named after the model or imply endorsement.
KimiMoonshot AIModified MIT on the K2 series: MIT, plus a requirement to display the model name in the product UI above 100 million monthly active users or 20 million US dollars in monthly revenue.The attribution clause is irrelevant at small scale and a real obligation at large scale. Later releases can carry different terms.
MistralMistral AI (France)Apache 2.0 for most recent open-weight releases, including the Mistral 3 family.Exceptions exist: some specialist models ship weights under non-commercial terms, and Mistral’s API-only models are not open at all.

Licence patterns are summarized from the vendors’ own licence texts and announcements. They are not legal advice.

Serving stack options

The serving engine decides how many users one GPU can carry. The choice is mostly about concurrency, hardware and how much you value a mature OpenAI-compatible API.

vLLM

The default for shared GPU serving. PagedAttention and continuous batching keep the card busy, the server speaks the OpenAI API, it supports tensor, pipeline, expert and data parallelism, prefix caching, and quantization formats including FP8, AWQ and GPTQ, and it exposes Prometheus metrics at /metrics.

SGLang

Worth measuring when your workload shares long prefixes, such as multi-turn agents or retrieval with stable system prompts. Hugging Face names it alongside vLLM as the recommended alternative to TGI.

Text Generation Inference (TGI)

Hugging Face has put TGI in maintenance mode, accepting only minor fixes, and recommends vLLM or SGLang for new endpoints. Existing TGI deployments still run; new builds should start elsewhere.

Ollama and llama.cpp

Excellent for development, workstations, edge and single-team tools. Both are built around GGUF and low concurrency; Ollama queues requests per model unless you tune parallelism. Move to vLLM or SGLang when many users share one GPU.

Which stack Swfte runs for you

Serving-stack choice is made per deployment, against the model, the hardware and the traffic shape, and recorded in the deployment’s release record. <serving stacks supported in Swfte-managed deployments — founder to fill>.

If you already have a tuned vLLM or SGLang deployment, bring it. Connect fronts any OpenAI-compatible endpoint, so the governed path (routing, approved-model policy, cost tracking, logging) does not depend on which engine sits behind it.

Quantization: memory you win and quality you must re-measure

Weights cost roughly two bytes per parameter in bf16, one byte in FP8 and about half a byte at 4-bit. That is plain arithmetic: a 70-billion-parameter model is about 140 GB of weights in bf16, about 70 GB in FP8 and about 35 GB at 4-bit, before KV cache and activations. An 80 GB H100 therefore cannot hold it in bf16 on one card, and can only hold the FP8 version with very little room for batching. Our GPU reference lists per-card memory so you can do this sum for your own model.

Mixture-of-experts models are the common trap. Only a fraction of parameters is active per token, which cuts compute, but every expert still has to sit in memory because the router picks different ones for each token. Active-parameter count is a speed claim, not a memory claim.

Quantization changes the model. The evaluation you ran on the bf16 checkpoint says nothing certain about the FP8 or 4-bit artifact you actually serve, particularly on refusal behavior and non-English text. Treat each quantized artifact as a new candidate and put it through the same gate.

Autoscaling and routing through the Connect gateway

Connect is Swfte’s universal model gateway: one API in front of many providers, with smart routing, automatic failover and cost tracking. For a private deployment it plays two roles. It is the single place where the approved-model policy is enforced on every request, and it is how you move traffic between an open-weight model on your own GPUs and a fallback without rewriting agents.

Scale on the signals that actually saturate an inference server: queue depth, time to first token and KV-cache utilization, not CPU. Keep a floor of warm replicas for latency-sensitive routes, because loading hundreds of gigabytes of weights is slow. Route low-risk, high-volume work to the smaller or cheaper model and keep the harder slice on the stronger one; the routing rule is policy, so it can be reviewed like any other change.

Routing is also your rollback lever. When the old and new revision are both deployed, shifting weight between them is a configuration change rather than a redeploy.

Evaluation gate, rollback and audit trail

Evaluation gate before promotion

A candidate revision is promoted only when it passes capability, safety, multilingual and regression suites against the version in production. A failed gate blocks the release; it is not a warning.

Rollback

Keep the previous approved revision deployed and routable until the new one has run on live traffic. Rollback is a routing change plus a log entry.

Audit trail

Record which model revision, weights hash, quantization, serving image and policy set served each request class, and who approved each promotion. That is the evidence you need when someone asks what ran.

Runtime guardrails

Input and output checks, tool-permission limits and human-approval rules sit around the model at request time. They answer should this be blocked; governance answers who was allowed to do what, and can you prove it.

Where it runs: dedicated, in-region, yours

Models run on dedicated infrastructure: single-tenant, in a region you choose, with options from an isolated VPC to bare metal in your own data center. Data residency options and private, on-premises and hybrid deployment are part of the platform’s stated direction, and the exact regions and facilities available to you are agreed per engagement. <EU regions and facilities offered — founder to fill>.

Swfte provides the technical controls, governance mechanisms and evidence you need to deploy AI within your applicable regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration.

Swfte does not hold a SOC 2 report or an ISO 27001 certificate and does not sign HIPAA BAAs today; a SOC 2 Type I audit is in preparation, and the trust page has the current status. Your own regulatory position depends on your use case, jurisdiction and configuration.

Deployment guides

Five guides, each for a different question.

Deploy an open-source LLM

The end-to-end path: sizing, engine, quantization, API surface and the evaluation gate.

Deploy an LLM in the EU

Residency versus sovereignty, transfer rules, EU AI Act timing and EU-language coverage.

Private LLM hosting

Single-tenant hosting patterns compared, and what privacy actually requires.

Self-hosted LLM inference

Engines, batching, KV cache, parallelism, metrics and a cost method.

Deploy an LLM on-premise

Air-gapped and on-premises: offline weights, verified bundles and patching.

Swfte Safety, our safety-first model

Design intent, model card and the fields still to be published.

Frequently asked questions

Which open-weight models can I deploy with Swfte?

Any open-weight model your licence review clears and that fits your hardware. Llama, Qwen, DeepSeek, Gemma, Kimi and Mistral families are the common starting points. Bring your own model, including your own fine-tune, goes through the same Pick, Evaluate, Harden, Deploy, Monitor loop.

Is deploying an open-weight model in the EU enough to keep data in the EU?

Hosting location is one part of it. Data residency covers where data is stored and processed; sovereignty also covers who can access it, which legal jurisdiction the operator is under, and which dependencies can reach it. Our EU deployment guide separates the two.

Which serving engine should I use?

For shared GPU serving, start with vLLM, and measure SGLang if your traffic shares long prefixes. Use Ollama or llama.cpp for development, edge and low-concurrency tools. Hugging Face has put TGI in maintenance mode and recommends vLLM or SGLang for new endpoints.

What does the evaluation gate check?

Capability on your own tasks, safety (refusal, over-refusal, jailbreak and prompt-injection resilience, harmful content, bias), multilingual behavior in EU languages, latency and cost, licence, and a supply-chain check on the weights. A candidate that fails does not get promoted.

Does quantizing a model change its safety behavior?

It can. Quantization changes the numbers the model computes with, so refusal behavior and non-English quality can shift. Re-run the evaluation suite on the quantized artifact you will actually serve.

Does Swfte hold security attestations?

Swfte does not hold a SOC 2 report or an ISO 27001 certificate and does not sign HIPAA BAAs today; a SOC 2 Type I audit is in preparation, and the trust page has the current status. Your own regulatory position depends on your use case, jurisdiction and configuration.

Does this settle my EU AI Act or GDPR position?

No tool can do that on its own. Swfte provides the technical controls, governance mechanisms and evidence you need to deploy AI within your applicable regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration.

Deploy your first model on Swfte

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.