What it actually costs to run a model locally
Not “best GPU” lists. For each model: the memory it really needs at each quantisation, which cards fit it, and what the same job costs to rent by the hour — computed from the model's own config and live provider pricing.
Interactive
Watch context length eat your GPU
Drag a slider and see the KV cache — not the weights — push a model off the card.
Live · hourly
Cloud GPU prices right now
Cheapest H100 and 5090 rental this hour, and the cheapest GPU that runs each model.
What fits at each VRAM tier
How many of the 40 tracked models run at a usable quant on each card class, at 16k context.
RTX 4060 · 3050
RTX 3060 · 4070
RTX 5060 Ti · 4060 Ti
RTX 3090 · 4090
2× 3090 · RTX 6000
On a 24 GB card at 16k context
Hardware guides
The longer decisions — which card, rent or buy, what the marketing leaves out.
GLM-4.5-Air on local hardware
The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.
Cheapest way to run GLM in the cloud
Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.
Run Mistral Small 24B locally
VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.
Cheapest RTX 4090 rental
Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.
Cheapest A100 80GB rental
Live A100 80GB prices cheapest-first, when the A100 beats an H100 on price-per-job, and the 70B-class workloads it's the sweet spot for.
Cheapest H100 rental
Live H100 prices across Vast and Runpod with the capture date, SXM vs PCIe, and the per-job maths for 70B–120B models.
H200 rental cost
141 GB is capacity, not speed — when the H200 premium over an H100 is worth paying, with live and published rates side by side.
RTX 5090 vs two RTX 3090s
Two 3090s give you 48 GB but not double the speed for chat — llama.cpp's default multi-GPU mode makes the cards take turns. What the benchmarks actually show.
Cheapest way to serve Llama 70B
"Cheapest" flips on duty cycle and concurrency — and in Europe the electricity bill alone can approach the cost of just renting. Plus the config defaults that silently break a 16k deployment.
What --n-cpu-moe actually does
The trick that makes MoE models 5× faster — except it usually makes them slower, and the famous speedup only happens when the model didn't fit in the first place.
Is the DGX Spark worth it?
NVIDIA spent three months optimising it and token generation got slower — because 128 GB of memory on a 273 GB/s bus holds huge models and reads them slowly. Who should actually buy one.
Memory figures are computed from each model's own config.json, using bits-per-weight constants checked against real published quantised file sizes (median error 0.7%). Rental prices are pulled hourly from provider APIs. Full methodology is on each model's page.