Confidential Computing for LLMs: What It Protects and Costs
Confidential computing protects LLM prompts and weights in use. How GPU TEEs work, their cost and their limits.
Confidential computing protects data while it is being processed, which is the one state that encryption at rest and in transit leaves exposed. For an LLM, it means the prompt, the retrieved documents and the model weights sit in memory that the host operator, the hypervisor and the cloud administrator cannot read, and a remote party can verify that before sending anything. It narrows who you have to trust. It does not make an AI system safe, correct or compliant by itself, and it has a measurable performance cost.
This explainer covers how confidential computing works for GPU inference, what independent benchmarks say about the overhead, what it protects against and what it does not, and where it fits next to on-premises, air-gapped and sovereign designs.
What problem does confidential computing solve for LLMs?
Data is normally protected in two states: encrypted on disk and encrypted on the wire. To compute on it, the system has to decrypt it into memory, and for that moment anyone with privileged access to the machine could in principle read it. In a cloud, privileged access includes the provider's hypervisor and administrators. When an LLM serves your prompts, that memory holds your most sensitive text, plus the proprietary model weights if you are a model owner.
Confidential computing uses hardware to create a trusted execution environment, or TEE, where memory is encrypted and isolated from the rest of the host. The three things that matter in practice:
- Isolation. The workload runs in an encrypted region the host cannot inspect.
- Attestation. The hardware produces signed evidence about what is running and on which firmware, so you can verify it before releasing data or keys.
- Key release tied to attestation. Secrets, such as the key that decrypts model weights, are released only to an environment that passes verification.
How do GPU TEEs work for inference?
LLM inference runs on GPUs, so a TEE that covers only the CPU is not enough. A typical setup, as described in the research literature, uses a confidential virtual machine on the CPU, with Intel TDX or AMD SEV-SNP, and an NVIDIA GPU such as the H100 operating in confidential computing mode. The CPU side provides the trusted root, and data moving between CPU and GPU is encrypted. A client can check the GPU's attestation report against NVIDIA's attestation service, for example to confirm the GPU is not running revoked firmware.
Cloud availability exists. Microsoft announced general availability of Azure confidential VMs with NVIDIA H100 GPUs, and a third-party comparison reports that Google Cloud also offers confidential GPU configurations on a different CPU technology. Check each provider's current documentation, since regions, GPU generations and instance sizes change.
What does confidential computing cost in performance?
Be skeptical of any single number. Published measurements disagree, and the differences are instructive.
- An independent benchmark of confidential GPU inference on H100 under Intel TDX reports time-to-first-token and request latency overheads from about 21 percent for a 7-billion-parameter model to roughly 27 to 30 percent for a 30-billion-parameter mixture-of-experts model, with global token throughput reductions of about 18 and 21 percent respectively. The authors recommend model-specific capacity planning rather than a fixed overhead assumption.
- Intel's own system-level analysis finds that smaller or launch-intensive workloads can see measurable degradation, while large, compute-dominated inference sees modest and predictable overhead.
- Some cloud and vendor guides state overhead of a few percent for large models. That is plausible for large, compute-bound workloads, but it is not what the independent H100 study measured on the models it tested.
- A study of newer Blackwell hardware reports low single-digit throughput overhead when the stack is configured correctly, and penalties of 30 to 40 percent for stock configurations that carry avoidable settings.
The sources of overhead are the encrypted transfers between CPU and GPU and the transitions in the confidential VM. Attestation itself is a one-time cost at startup in one vendor's description, on the order of seconds, not a per-request tax.
The planning conclusion: expect a real cost, measure it on your model, batch size and context length, and treat any vendor figure as a hypothesis. Our posts on continuous batching and serving frameworks show why batching changes the picture.
What does a TEE protect against, and what does it not?
| Threat | Does confidential computing help? |
|---|---|
| Cloud administrator or hypervisor reading memory | Yes, this is the core purpose |
| Other tenants on shared hardware | Yes, through isolation |
| Stolen disk or snapshot | Yes, with encryption plus key release tied to attestation |
| Model weights copied by the host operator | Yes, when weights are decrypted only inside the TEE |
| Prompt injection or jailbreaks | No |
| An application that logs prompts in plaintext | No, it protects against the host and not against your own code |
| Over-broad retrieval that exposes documents to the wrong user | No |
| A malicious or buggy model server image | Only if you attest the exact code and measure it |
| Side-channel and physical attacks on the hardware | Not fully; research on TEE side channels is ongoing |
| Compelled disclosure by law of the provider | Not a substitute for jurisdictional analysis |
Two lessons from that table. First, attestation is only as good as what you verify: if you attest "a confidential VM" and not "this exact image," you have proven little. Second, confidential computing and governance are complementary. Governance decides what the AI may do, and the TEE limits who can observe the runtime.
How does it relate to sovereignty and private AI?
Confidential computing changes the trust model for a hosted environment: you can rely on hardware-enforced isolation and verification rather than only on the provider's operational promises. That is valuable when you use a third party's infrastructure for sensitive workloads. It does not change the provider's legal jurisdiction. If your question is "which law can compel the operator," read sovereign AI versus private cloud. If your requirement is that data never leaves your premises, a TEE in someone else's data center is the wrong tool, and you want enclosed AI or an air-gapped deployment.
A rough decision guide:
- Data may be processed by a cloud under contract, but you want the operator blind to it: confidential computing is a strong fit.
- Data must not leave your premises: on-premises or air-gapped, with confidential computing optional as defense in depth.
- Mixed estate: route by data class through a gateway, and enforce attestation for the confidential tier.
A checklist before you adopt it
- Define the threat: which actors do you want blind to the data?
- Confirm the full chain: CPU TEE, GPU mode, firmware and driver versions all attested.
- Decide what you attest: the exact container image or boot measurement, not just the platform.
- Tie secret release to attestation, including weight decryption keys.
- Benchmark your model at your concurrency and context length, in both modes.
- Plan for operations: debugging is harder when you cannot inspect the host.
- Keep the application layer honest: logging, retrieval permissions and agent policy still need design. The AI DMZ post covers controls around the model.
- Document the residual trust: the hardware vendor, the firmware and your own software.
Where does Swfte fit?
Swfte approaches protection in layers. Cortex keeps sensitive work on the device by default, dedicated cloud provides isolated infrastructure, BuildX routes by policy across models, and Studio and Nexus govern what agents can do. Hardware-attested confidential computing is an option to evaluate for hosted environments where it matches the threat model; the infrastructure and governance pages describe the layers.
Swfte is designed to provide technical controls, governance mechanisms and evidence for deploying AI within your own regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration. To discuss your threat model, contact the team. For a self-hosted view, see the self-hosted LLM stack guide.
Frequently asked questions
What is confidential computing for LLMs?
It is the use of hardware trusted execution environments, on both CPU and GPU, to keep prompts, documents and model weights encrypted and isolated while an LLM processes them, with remote attestation so you can verify the environment first.
Does confidential computing slow down inference?
Yes, by an amount that depends on the model, batch size and configuration. An independent H100 study measured roughly 12 to 30 percent depending on the metric and model, while other sources report a few percent for large models. Benchmark your own workload.
Can the cloud provider read my prompts in a confidential VM?
The design goal is that the operator and hypervisor cannot read TEE memory. That depends on correct configuration and on verifying attestation. It does not address legal compulsion of the provider.
Is confidential computing the same as air-gapped?
No. Air-gapping removes the network path. Confidential computing keeps the network but protects data in use from the host. They solve different problems and can be combined.
Which GPUs support it?
NVIDIA's H100 supports confidential computing mode, and newer Blackwell parts have been benchmarked. Availability by provider and region changes, so confirm with current documentation.