Deployment guide

Private LLM hosting: single-tenant models behind your own controls

What private actually requires, the hosting patterns compared, and the controls that make it hold.

Private LLM hosting means your prompts, your documents and your outputs are processed by a model you control, on infrastructure no one else shares, with access you can audit. The word “private” gets applied to everything from a vendor’s enterprise API tier to an air-gapped rack, so the useful question is not whether a service is private but which specific exposures it removes.

Five hosting patterns, compared honestly

PatternWho runs the modelWhat it removesWhat remains
Shared model APIThe model lab, multi-tenantNothing structural; contractual terms govern use.Prompts leave your environment; you depend on the lab’s retention, region and policy changes.
Enterprise API tier or dedicated endpoint at the labThe model lab, isolated capacityNoisy neighbors; some retention terms.Prompts still leave your environment and the lab is still in the loop.
Open-weight model on single-tenant cloudYou, or Swfte for you, on dedicated infrastructure in a region you chooseThe lab-side API, shared hardware and region ambiguity.Cloud operator and its jurisdiction; your own patching and operations.
Open-weight model on-premisesYou, in your data centerThird-party operator access to the hardware.Procurement, capacity planning, physical security and operations are yours.
Air-gappedYou, with no outbound networkNetwork-borne exfiltration paths.Update logistics, offline tooling and the discipline to keep it truly isolated.

What “private” actually requires

Isolation is necessary but not sufficient. The model server should sit on a private network with no public ingress, reached only by the gateway that enforces policy. Authentication belongs on that gateway and on the engine itself; some engines authenticate only certain API paths, so check what your engine protects and put the rest behind the network boundary.

Then look at the data you are creating. Prompts and completions are logs the moment you record them. Decide what you log, redact personal and confidential content before it reaches the log store, set a retention period, and keep the logs in the same region as the model. Keys, including the keys that decrypt weights at rest and the credentials the gateway holds, belong in a key-management service with rotation and access control, not in environment files.

Finally, bound what the model can do. A private model that can call tools can still be talked into misusing them. Prompt injection, direct or through retrieved documents, is the top-ranked risk in the OWASP Top 10 for LLM Applications. Treat retrieved text as data, scope tool permissions to the least needed, and require human approval for consequential actions.

A private-hosting control checklist

  • Dedicated, single-tenant compute; no shared GPUs between customers.
  • Private network only; the gateway is the only client of the model endpoint.
  • Weights verified against pinned hashes and loaded from safetensors; remote code execution switched off.
  • Authentication on the gateway and the engine; per-workload credentials with rotation.
  • Encryption in transit and at rest, with customer-controlled keys where the deployment allows it.
  • Prompt and completion logging with redaction, retention and in-region storage.
  • Approved-model policy enforced per workload, so a team cannot quietly point at a different model.
  • Runtime guardrails and human-approval rules for consequential tool calls.
  • Evaluation gate on every model, quantization and prompt change, with rollback to the last approved revision.
  • An audit trail linking each request class to the model revision, serving image and policy set that produced it.

Private LLM hosting with Swfte

Swfte runs open-weight and private models on dedicated infrastructure, from an isolated VPC to bare metal in your own data center, and puts them behind Connect so approved-model policy, routing, failover and cost tracking apply to every request. Nexus adds identity, permissions and an audit trail for the agents that call them. Private, on-premises and hybrid deployment are part of the platform’s stated direction; the exact options for your environment are agreed per engagement.

Swfte does not hold a SOC 2 report or an ISO 27001 certificate and does not sign HIPAA BAAs today; a SOC 2 Type I audit is in preparation, and the trust page has the current status. Your own regulatory position depends on your use case, jurisdiction and configuration.

Frequently asked questions

What is private LLM hosting?

Running a language model on infrastructure you control or that is dedicated to you, so prompts and outputs are processed without a shared multi-tenant model API in the loop, with access you can audit.

Is an enterprise API tier from a model lab private?

It can reduce noisy-neighbor and retention exposure, but prompts still leave your environment and the lab remains in the loop. A model you host yourself removes the lab-side API entirely.

Does a private model still need guardrails?

Yes. Privacy controls where data go; guardrails and governance control what the model can do. Prompt injection and excessive agency are risks regardless of where the model runs.

Do I have to log prompts?

Not necessarily, but logs are what make an audit trail possible. If you log, redact sensitive content, set retention and keep logs in-region.

Does Swfte hold security attestations?

Swfte does not hold a SOC 2 report or an ISO 27001 certificate and does not sign HIPAA BAAs today; a SOC 2 Type I audit is in preparation, and the trust page has the current status. Your own regulatory position depends on your use case, jurisdiction and configuration.

Host a private model with Swfte

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.