What it costs to run Gemma 2 27B
Trying to decide what to buy? The best GPU for Gemma 2 27B →
We don't yet have the layer/attention shape for this model — its repository is licence-gated, so the KV-cache figures below can't be computed. Weight sizes are still accurate.
How these numbers are produced
- Model shape comes from the model's own config.json.
- Weight size is parameters × effective bits-per-weight. Those constants are checked against real published quantised file sizes — median error 0.7% across 40 measurements.
- KV cache is 2 × layers × KV-heads × head-dim × context × 2 bytes. Using KV-heads rather than attention heads is what makes this correct for grouped-query attention; treating a GQA model as multi-head overstates the cache by up to 8×.
- Tokens/sec is an estimate, not a benchmark. Generation is memory-bandwidth-bound, so this is bandwidth ÷ weight-bytes derated to 75%. Real throughput depends on your runtime, batch size and kernels. We don't run our own hardware tests — see our editorial policy.
- Rental prices are pulled hourly from provider APIs (last updated Sat, 10 Oct 2026 11:21:09 GMT). Spot/interruptible pricing can change or vanish without notice.
Try it interactively
Drag the context slider and watch the KV cache fill the card
Worth reading before you buy
Cheapest RTX 4090 rental
Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.
GLM-4.5-Air on local hardware
The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.
Cheapest way to run GLM in the cloud
Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.
Run Mistral Small 24B locally
VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.