Swfte Connect / Self-deploy

Run the LLM gateway in your own cloud, VPC or data centre

Swfte Connect is the model gateway and routing layer of the Sovereign Intelligence Platform. You can use it as a managed service, or deploy it into infrastructure you control, so prompts, model traffic and keys stay inside your boundary.

Why teams run the gateway themselves

A gateway sits in the path of every prompt and every response. It sees the data, holds the provider keys and decides which model answers. That position is exactly why some organizations want it inside their own network.

Three reasons come up most. First, data location: when the gateway and the inference runtime run in your VPC, traffic to your own models never has to cross the public internet or leave your cloud account. Second, key custody: provider keys can live in your secret store and be used only from your network. Third, independence: a gateway you operate keeps serving your self-hosted models even if the managed service is unreachable.

None of this removes the reasons to use the managed service. If you want to start in minutes, avoid operating infrastructure and route to hosted providers, the managed Connect is the quicker path. Self-deployment is for teams whose constraints, not preferences, decide where AI runs. The comparison table later on this page lays the two modes side by side.

Deployment forms

Connect is designed to run in the forms below. The status column is deliberate: it separates what the tooling already generates from what is arranged with the team.

Connect deployment forms and their status
FormWhat it isStatus
Managed (SaaS)Connect operated by Swfte. Customer data is stored in AWS eu-west-1 (Ireland) today.Available now
Your cloud account or VPCTerraform generated for AWS, GCP and Azure, deploying the gateway proxy and model runtime into your account.Generated by Swfte tooling, provided on request
KubernetesA Helm chart for the same components, for clusters you already operate.Generated by Swfte tooling, provided on request
Single host or on-premDocker Compose output for a server or a small cluster in your own data centre.Generated by Swfte tooling, provided on request
Air-gappedOperation with no outbound path to Swfte, using an offline-capable license and your own model weights.Designed for, on request
Dedicated, Swfte-operatedPrivate capacity run for you on the dedicated cloud, with data location agreed per deployment.On request, see Dedicated cloud

The generated bundles cover the model infrastructure (an inference runtime and a gateway proxy), and can also include an agent runtime, a workflow executor and tool services, so the same environment can later host more than the gateway. A license key is required to generate a bundle.

Architecture

A self-deployed Connect is a small number of parts, all running inside your boundary.

  1. Your applications and agents call one OpenAI-compatible endpoint. Moving an existing integration over is a base-URL and key change.
  2. The gateway proxy authenticates each call, applies rate limits, meters usage and forwards the request. It is deliberately a thin layer, so it stays easy to review and to update.
  3. The inference runtime serves open-weight models on your GPUs. The reference runtime is vLLM, run as the official container image.
  4. Optional external providers are reached through your own egress, using keys you hold, only for the models and providers you choose to enable.
  5. Optional agent runtime, workflow executor and tool services run alongside, so agents and workflows can use the gateway without leaving the environment.

Scope note: the generated bundle runs the gateway proxy in front of your own inference. Running the complete multi-provider routing, budget and fallback layer of the managed Connect inside your VPC is designed for and arranged on request. Ask the team which components your target version includes.

What stays in your environment, and what can leave

The point of self-deployment is a clear answer to "where does this go?". Here it is, item by item.

Data and control boundaries in a self-deployed Connect
ItemWhere it livesNotes
Prompts and responses to self-hosted modelsYour environmentThe request goes from the gateway proxy to your inference runtime and back.
Model weightsYour environmentStored and loaded from storage you control.
Provider API keys (BYOK)Your secret storeUsed from your network when you enable an external provider.
Requests to external providersLeave through your egressOnly for providers you enable. The provider then processes them under its own terms.
Usage telemetryOptionalA configuration switch controls whether usage telemetry is sent to Swfte.
License validationSigned key, offline graceKeys are signed and validated locally, with a grace period if the license server cannot be reached.

Configuration

Standing up a deployment follows the same order each time.

  1. Choose the target. Pick AWS, GCP, Azure, an existing Kubernetes cluster or a Docker host, and the region or site. Because the output is Terraform, Helm or Compose, it reads as code and goes through your normal review.
  2. Provide the license key. It is signed and checked by the gateway proxy.
  3. Select models. Choose open-weight models from the catalogue to host, and note the GPU memory each needs.
  4. Set authentication and limits. Issue API keys, set rate limits, and where your license includes it, connect single sign-on.
  5. Add external providers if you want them. Supply your own keys; they stay in your secret store.
  6. Set telemetry. Decide whether usage telemetry is sent, and keep logs where your own tooling collects them.
  7. Point your applications at the new endpoint, keeping the same request format.

Upgrades and operations

Operating a gateway means you own the change window. The tooling is built so that a change is a reviewable diff, not a console session.

Because the deployment is generated as infrastructure code, an upgrade is a new set of artifacts applied through your own pipeline. You see what changes before it applies, you can roll it out to a staging environment first, and you can roll back by re-applying the previous version. Release cadence, version-support policy and the image registry are set out in the onboarding material: <upgrade cadence and version policy - founder to fill>.

Licensing is designed not to be a point of failure. The license is a signed key with a grace period, so a temporary loss of connectivity does not stop the gateway. The design specification sets the grace period at 30 days, and the license model includes an air-gap capability for sites with no outbound path.

Governance in a self-deployed gateway

Controls that matter for AI traffic come with the gateway, wherever it runs.

  • Identity and access. Calls are authenticated; permissions are scoped to workspaces; the license model includes single sign-on.
  • Policy. Content policies can detect secrets and personal data in requests and redact them before the request goes to a model.
  • Budgets. Usage caps for tokens and spend, per workspace or per model, with automatic downgrade to a cheaper model when a cap is reached, plus email and webhook alerts.
  • Audit and evidence. Usage is metered and action audit events are recorded per workspace, so a reviewer can start from an export.

These are the same controls described on the security page and map to the Trust and Governance Fabric of the platform. Self-deploying changes where they run, not whether they exist.

Managed Connect compared with self-deployed Connect

Managed and self-deployed Connect compared
DimensionManaged (SaaS)Self-deployed
Where traffic runsSwfte-operated, AWS eu-west-1 (Ireland) todayYour cloud account, VPC or data centre
Who operates itSwfteYou, with Swfte tooling and support
Time to first requestMinutes, through self-serve sign-upA deployment project, planned with the team
Provider accessPooled access or your own keysYour own keys, through your egress
Hosted open-weight modelsThrough Swfte inference capacityOn GPUs you provide
UpdatesApplied by SwfteApplied by you, through your pipeline
TelemetryPart of the serviceSwitchable
Best forFast start, no infrastructure to runStrict data location, network isolation, own-cloud commitments

Who should self-deploy

Self-deploy if your policy requires prompts to stay inside a boundary you control, if you hold committed capacity or reserved GPUs in your own cloud, if the workload runs in a network with no direct internet path, or if you need the gateway to keep working independently of any outside service. Stay on the managed service if speed matters most and the data you send is already cleared for a hosted provider. Many teams do both: managed for experiments, self-deployed for the workloads that need it, with one API shape across both.

This page describes capabilities and deployment intent, not a certification of any customer deployment. Swfte provides the technical controls, governance mechanisms and evidence required to deploy AI within an organization's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration.

Frequently asked questions

Can I run Swfte Connect in my own VPC?

Yes, that is the design. Swfte generates Terraform for AWS, GCP and Azure, plus Helm and Docker Compose output, that deploy a gateway proxy and an inference runtime into your own environment. The artifacts are provided on request, and a license key is required.

Does self-deployed Connect need to call Swfte?

Not for your model traffic. Prompts to your self-hosted models stay inside your environment. Usage telemetry is a configuration switch, and licensing uses a signed key with a grace period, with an air-gap capability for sites with no outbound path.

Is air-gapped deployment available?

It is designed for and arranged on request. The license model includes an air-gap capability and offline grace, and you supply the model weights. Talk to the team about your site and constraints.

Which models can a self-deployed Connect serve?

Open-weight models you host on your own GPUs, through a vLLM runtime, and optionally external providers using your own keys. See the providers page for the families Connect routes to.

How do upgrades work?

The deployment is code, so an upgrade is a new set of Terraform, Helm or Compose artifacts that you review and apply through your own pipeline, staging first. <upgrade cadence and version policy - founder to fill>.

Is a self-deployed gateway compliant by itself?

No deployment form is compliant by itself. Self-deploying gives you the controls and the evidence, under your own configuration. Swfte provides the technical controls, governance mechanisms and evidence required to deploy AI within an organization's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration.

Connect in the platform

Route every model through one governed gateway

Start managed in minutes, or plan a deployment inside your own boundary with the team.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.