What a GPU run actually costs
The arithmetic is simple and the judgement is not:
cost = rate ($/GPU-hour) × GPUs × hours × runs
Worked example: 8 GPUs at $2.50/hour for 12 hours, four runs a month:
- Per run: 2.50 × 8 × 12 = $240
- Monthly: 240 × 4 = $960
- Yearly: $11,520
- GPU-hours consumed: 8 × 12 × 4 = 384 GPU-hours
The number people forget is the last one. 384 GPU-hours is 384 hours of a machine you did not have — which is why the rental-versus-purchase comparison is not close for occasional work and reverses quickly for sustained use.
Renting versus owning
Break-even is straightforward: purchase price ÷ (hourly rate × GPUs) GPU-hours. At $2.50/hour on 8 GPUs, a $20,000 GPU breaks even after 1,000 GPU-hours, which is 125 runs or about 10 months of this workload.
That arithmetic says nothing about the four things that usually decide it:
- Utilisation. A GPU used 10% of the time is 90% wasted capital. Owning makes sense above roughly 40% sustained use for a single job, and considerably higher if your workload is bursty and you would otherwise queue for capacity.
- Depreciation and obsolescence. Consumer hardware loses value fast, and a five-year-old training GPU is not what you want for a new architecture. A five-year straight-line depreciation is optimistic for consumer cards; three years is more realistic.
- Power and space. A 700 W GPU plus the rest of the system is around 1 kW, which is $30-60 a month in electricity depending on where you are, plus the space and the cooling. This narrows the gap but rarely closes it.
- Opportunity cost. Money in a GPU is money not in people or in serving capacity. For most teams the honest comparison is renting until the workload is reliably large.
Instance types change the number by 6x
- On-demand — guaranteed, most expensive, the right default for anything with a deadline.
- Spot / preemptible — 50-80% cheaper, reclaimed with short notice. Suits training that checkpoints frequently, batch evaluation, and anything restartable. Not suitable for a single long run that cannot resume.
- Reserved / committed — 30-60% cheaper than on-demand for a one-to-three-year commitment. Worth it only if the utilisation is genuinely sustained; a reservation you stop using is more expensive than on-demand.
The alternatives are usually cheaper than they look
Before renting a GPU for inference, check whether you need one at all. For many serving workloads the options are considerably cheaper per unit of output:
- API access for inference. You pay per token with no idle cost, and at modest volume it is far cheaper than running your own GPU. A single mid-range GPU serving a small model can cost more per month than the API bill for a surprisingly large user base.
- Serverless inference. Scale to zero between bursts; you pay only for the requests you serve.
- Quantisation. Running a model at 4-bit or 8-bit often halves the GPU memory needed, which frequently means a smaller and much cheaper instance — or the same instance serving two to four times the concurrency.
- Smaller models. A well-chosen small model can match a large one on a narrow task at a fraction of the cost. This is the routing argument from the AI API cost guide applied to self-hosting.
For fine-tuning specifically
LoRA and QLoRA cut the cost of adapting a model by one to two orders of magnitude versus full fine-tuning, because the optimiser states shrink from billions of parameters to a fraction of a percent. A full fine-tune of a 7B model on a single GPU is a research project; the same adaptation with QLoRA on one consumer card is an afternoon.
The recurring cost then dominates: one training run is a fixed price, but serving the result is monthly. Budget for the second one, not the first.