Quick answer
What should you choose?
Vast.ai exposes marketplace offers, RunPod publishes self-service GPU and serverless paths, Thunder Compute lists direct per-minute instances, and Massed Compute offers hourly and dedicated models. The lowest total cost depends on required VRAM, completed-job time, storage, billing unit, interruption risk and whether the workload can scale to zero.
Hourly price is only one line in a GPU job's cost. An undersized card that cannot load the model, a preempted training run without checkpoints or an idle instance left running can erase an apparent saving.
Normalize every candidate to the same completed workload. Record setup time, actual run time, retry probability, storage retained after shutdown and the steps required to stop every billable resource.
Current facts that change the decision
Live listings vary by host, verification, region and interruptibility; compare the exact offer.
Choose based on utilization and control; include persistent storage separately.
Compare the live GPU rate plus selected CPU, RAM, disk and retained-snapshot cost; delete finished instances to stop compute billing.
Compare the break-even point between short jobs and continuously used capacity.
Time-sensitive facts verified August 19, 2026. Always recheck the live product page before paying.
The shortlist at a glance
Start with buyer fit, then validate the exact plan. Candidate order follows this guide's decision path; it is not a synthetic score.
Vast.ai
A distributed GPU marketplace with variable host, price and reliability characteristics.
Price-sensitive experiments that can compare individual marketplace offers
You need a uniform provider-wide hardware and support promise
RunPod
A US-operated AI cloud combining GPU Pods, serverless inference and clusters.
Developers moving between GPU development and production inference
You have not separated storage and idle-resource cost from compute
Thunder Compute
A US-operated GPU cloud with per-minute instances, retained snapshots and direct billing controls.
Buyers who want a direct North America GPU instance billed per minute
You need a managed model API, distributed marketplace or serverless inference product
Compare every candidate
| Provider | Best fit | Key limitation | Company region |
|---|---|---|---|
| Vast.aigpu-cloud | Price-sensitive experiments that can compare individual marketplace offers | You need a uniform provider-wide hardware and support promise | United States |
| RunPodgpu-cloud | Developers moving between GPU development and production inference | You have not separated storage and idle-resource cost from compute | United States |
| Thunder Computegpu-cloud | Buyers who want a direct North America GPU instance billed per minute | You need a managed model API, distributed marketplace or serverless inference product | United States |
| Massed Computegpu-cloud | Teams that may grow from one GPU into dedicated or clustered capacity | You only need a managed pay-per-token model API | United States |
How to choose without buying the wrong plan
- Reject GPUs that cannot fit the workload with headroom
- Estimate successful completed-job hours rather than requested hours
- Price checkpoints and persistent storage
- Model interruption and restart risk
- Create an automatic or documented teardown step
A current offer is not automatically the lowest total cost. Compare the initial charge, billing period, renewal amount, required add-ons, backups, migration effort and your administration time.
Frequently asked questions
How do I compare GPU cloud prices fairly?
Use the same GPU memory requirement, precision, workload, expected successful runtime, storage and transfer assumptions. Then compare total completed-job cost rather than hourly price alone.
Can a smaller GPU cost more?
Yes. If it forces a slower run, smaller batch, model offloading or repeated failures, a lower hourly rate can produce a higher completed-job cost.
When is serverless GPU cheaper?
It can be cheaper for bursty inference with long idle periods, but request latency, cold starts, minimum billing units and per-request overhead still need measurement.
Primary sources
- Vast.ai platform and live marketplace ↗
- RunPod Cloud GPU pricing ↗
- Thunder Compute pricing ↗
- Thunder Compute billing documentation ↗
- Massed Compute product catalog ↗
Recheck the exact plan, company terms and checkout total before buying. Product pages and availability can change after the verification date.



