GPU buying guide
Cloud-based GPU: four ways to rent one and how to compare them
Understand what a cloud GPU is, the four ways to get one, and the questions that decide cost and legal exposure before you sign.
A cloud-based GPU is a graphics processor in a provider’s data centre that you use over the network, as a virtual machine, a container, a per-request service or a dedicated cluster. You can get one from a hyperscaler instance, a GPU-specialist cloud, a serverless GPU or inference API, or a dedicated deployment. This page publishes no prices because they change daily. It shows how to read a listing, how to compare cost per useful token, and which questions about quotas, egress and jurisdiction to ask first. Swfte does not rent raw GPUs.
Last verified 2026-10-07. Sources are listed at the end of the page.
What are the four ways to get a cloud GPU?
Pick the way by how steady your load is and how much of the stack you want to run. The categories overlap at the edges.
| Way | What you rent | Suits | Watch for |
|---|---|---|---|
| Hyperscaler instances | A virtual machine with one or more GPUs from AWS, Google Cloud or Azure. | Teams already on that cloud, with identity, networking and procurement in place. | GPU quotas, regional availability and egress fees: read each provider’s documentation. |
| GPU-specialist clouds | Virtual machines or bare metal from providers that sell mainly accelerators. | Training and steady inference where the listing, interconnect and support fit. | The regions list, the legal terms and which managed services surround the GPU. |
| Serverless GPU and inference APIs | A container that scales with requests, or a hosted model behind an API you call per token. | Spiky or low-volume traffic, and teams that do not want to run drivers and engines. | Ask about cold starts, per-request limits, and how much control you have over the model build and the data path. |
| Dedicated or private deployment | Capacity reserved for you, in a region and under terms you agree, often in your own cloud account. | Regulated workloads and steady high volume with data-location requirements. | Ask about lead time, minimum terms, and the operations work you take on or contract out. |
How do I read a GPU instance listing?
Every provider lists the same few facts. These examples come from the providers’ own pages and NVIDIA’s, read on 2026-10-07. They show what to look for, and are not a recommendation.
| Example listing | What the page states | What it tells you |
|---|---|---|
| AWS P5 instances | NVIDIA H100 Tensor Core GPUs, up to 8 per instance, up to 640 GB of HBM3 GPU memory per instance. | Memory per instance, not per card. Divide by the GPU count before sizing a model. |
| Azure ND H100 v5 | Eight NVIDIA H100 GPUs (80 GB each in the spec table), NVLink inside the VM and InfiniBand between VMs. | Interconnect matters once a model spans more than one GPU or node. |
| Azure ND MI300X v5 | Eight AMD Instinct MI300X GPUs with 192 GB each, per the spec table. | Not every cloud GPU is NVIDIA, so check that your serving engine supports the vendor. |
| Google Cloud A4 and A3 Ultra | NVIDIA B200 on A4 with 1,440 GB of HBM3e across 8 GPUs, and H200 on A3 Ultra with 1,128 GB across 8 GPUs. | Machine series names map to a GPU model. The GPU docs page also links regional availability. |
| Google Cloud Run with GPUs | NVIDIA L4 with 24 GB, or RTX PRO 6000 Blackwell with 96 GB. Instances can scale down to zero. | A serverless form of the same chips. Quota is granted per project, and more is requested separately. |
| NVIDIA H100 data-centre page | 80 GB for the SXM form and 94 GB for the NVL form. | The same chip name comes in variants, so confirm which one the provider rents. |
Sources: AWS P5 page, Azure ND family page, Google Cloud GPU and Cloud Run GPU pages, and NVIDIA’s H100 page. Listings change as providers add GPUs.
How do I compare cost per useful token?
An hourly rate alone misleads. Compare providers on what a GPU produces for you, using arithmetic you can check and your own measurements.
1. Fit the model
Weights need roughly parameters times bytes per parameter in memory, plus KV cache and headroom. Pick the smallest GPU count that holds your model at your chosen precision and context length.
2. Measure sustained tokens per second
Serve the model with the engine you will use and replay real prompts at your target concurrency. Record the throughput that meets your latency limit, not the peak.
3. Apply real utilisation
A GPU idle at night still bills on an on-demand or reserved instance. Divide by the share of hours it does useful work.
4. Divide
Hourly cost divided by useful tokens per hour gives cost per token. Add storage, egress and the engineer time to run it, then compare with the per-token price of an inference API for the same model.
5. Re-run quarterly
Providers add GPU generations and change terms. Keep the spreadsheet, and refresh it from each provider’s current pricing page.
What should I ask a provider before committing?
- Quotas. Does your account start with a GPU quota, and how long does an increase take? AWS documents per-Region quotas measured in vCPUs for On-Demand instances and a console request process. Google documents an initial Cloud Run GPU quota and a request route for more.
- Reserved versus on-demand. AWS Capacity Blocks for ML reserve GPU instances for a future date for a fixed period, up to eight weeks ahead, and the page says cancellations are not allowed. Instances in a block do not count against On-Demand limits. Other providers have equivalents, so ask what you can cancel.
- Egress and switching. What does moving data out cost today? The European Commission states that the Data Act removes switching charges, including data egress, from 12 January 2027, with providers allowed to charge for them until then. Not verified here: whether any later amendment changes that date.
- Region and availability. Is the exact GPU offered in the region you need? Providers list this per GPU and per region, and it differs.
- Who can compel access. See the next section: ask which legal system governs the provider and its parent.
- Support and failure handling. What happens to a reserved block or a node when hardware fails, and who replaces it.
Does region alone decide who can access my data?
No. Where a GPU sits and which law reaches the company that runs it are different questions. The US CLOUD Act added 18 U.S.C. § 2713, which covers preservation and disclosure of communications and records “regardless of whether such communication, record, or other information is located within or outside of the United States.” The Department of Justice’s CLOUD Act page describes the Act as also authorising bilateral agreements under which foreign partners can obtain electronic data held by US-based providers.
Whether a particular provider and a particular order fall under that section is a legal question for your counsel. The practical step is to ask each provider where its parent company is established, what its policy is on government requests, and whether customer-managed keys or a dedicated tenancy are offered.
European buyers can borrow a structure instead of inventing one. The European Commission describes its Cloud Sovereignty Framework as a tool to evaluate providers of sovereign cloud, with 48 criteria in eight categories, and “legal and jurisdictional” is one of the eight. You can use the same categories as headings in your own provider questionnaire.
Where Swfte fits
Swfte is not a GPU cloud and does not rent raw GPUs here. If you only need a GPU, go straight to a provider above. You do not need Swfte for that.
Swfte helps with what runs on the GPU. Connect is one OpenAI-compatible API in front of hosted providers and your own endpoints, with routing, fallback chains, budgets and an audit event stream. Connect self-deploy generates Terraform, Helm and Docker Compose for AWS, Google Cloud and Azure, available on request and with a licence key. Dedicated deployment, including where it runs, is scoped through an engagement and is designed for, not self-serve. The GPU reference lists accelerator specs, and dedicated cloud explains the single-tenant option.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- AWS EC2 P5 instances. GPU model, GPUs per instance and memory per instance.
- AWS EC2 Capacity Blocks for ML. Reservation behaviour, limits and cancellation.
- AWS EC2 service quotas. Per-Region quotas and how to request an increase.
- Google Cloud GPU machine types. GPU models, machine series and memory.
- Google Cloud Run GPU support. Serverless GPUs, scale to zero and initial quota.
- Azure ND family VM sizes. GPU model, GPU memory and interconnect per series.
- Azure VM sizes overview. GPU-accelerated size families (NC, ND, NG, NV).
- NVIDIA H100 data-centre GPU. Memory for SXM and NVL variants.
- NVIDIA L4 data-centre GPU. L4 memory size.
- 18 U.S.C. § 2713 (Cornell LII). Statutory text on data located inside or outside the United States.
- US Department of Justice: CLOUD Act resources. The Department’s description of the Act and executive agreements.
- European Commission: Sovereign Cloud Framework explained. Framework scope: 48 criteria, eight categories.
- European Commission: Data Act explained. Cloud switching and egress charge dates.
Frequently asked questions
What is a cloud-based GPU?
A cloud-based GPU is a graphics processor hosted in a provider’s data centre that you rent over the network instead of buying. It reaches you as a virtual machine, a container that scales to zero, a hosted model priced per token, or a dedicated cluster. You pay for time or usage, and you can change size without buying hardware.
Is renting a cloud GPU cheaper than buying one?
It depends on utilisation, and this page states no prices. Renting avoids the capital outlay and the operations work, while owning can win when a GPU stays busy for years. Compute cost per useful token for both: hourly or amortised cost divided by sustained tokens per hour at your latency target, then add power, space, staff and egress.
How much GPU memory do I need for an LLM?
Weights need roughly parameters times bytes per parameter, so a model with 70 billion parameters at 8-bit precision needs about 70 GB for weights alone. KV cache and runtime overhead come on top, and grow with context length and concurrent users. Divide the provider’s memory figure by its GPU count to compare per-card memory.
Does the CLOUD Act apply to data stored in the EU?
The statute covers data held by in-scope providers regardless of whether it is located inside or outside the United States, so a European region does not settle the question by itself. Whether a given provider is in scope depends on its legal ties. Ask the provider, and take legal advice for your own case.
Does Swfte rent GPUs?
No. Swfte is not a GPU cloud and does not rent raw GPUs. It offers Connect, a gateway in front of models, and generators for Terraform, Helm and Docker Compose that deploy into your own AWS, Google Cloud or Azure account on request. A dedicated deployment is scoped through an engagement with the team.
When is a serverless GPU a better fit than an instance?
Serverless suits spiky or low-volume traffic, because the service can scale to zero when idle, as Google documents for Cloud Run GPUs. An instance suits steady load that keeps the GPU busy. Test cold-start time, per-request limits and model-size limits against your own workload before deciding.