What it costs to run Gemma 4 26B-A4B (MoE)
Trying to decide what to buy? The best GPU for Gemma 4 26B-A4B (MoE) →
Cheapest way to rent it right now
$0.05/hr
NVIDIA GeForce RTX 3060 12GB · Q2_K · spot · ≈ $13/mo at 8h/day
Live price captured Thu, 24 Sep 2026 21:20:58 GMT. Referral link — we may earn a commission at no cost to you; it never changes which card is cheapest.
Memory needed, by quantisation
At 16k context. Weights + KV cache + ~0.8 GB overhead.
| Quant | Weights | KV cache | Total | Quality |
|---|---|---|---|---|
| F16 | 49.4 GB | 0.3 GB | 50.5 GB | Full precision. Reference quality, twice the size of Q8 for no practical gain in most chat use. |
| Q8_0 | 26.3 GB | 0.3 GB | 27.4 GB | Effectively lossless. Use when VRAM is not the constraint. |
| Q6_K | 20.3 GB | 0.3 GB | 21.4 GB | Very close to Q8 at meaningfully less memory. A safe high-quality pick. |
| Q5_K_M | 17.6 GB | 0.3 GB | 18.6 GB | Small, hard-to-notice quality loss. Good balance. |
| Q5_0 | 17.1 GB | 0.3 GB | 18.2 GB | Older-style 5-bit. Q5_K_M is usually the better pick at the same size. |
| Q4_K_M | 15.0 GB | 0.3 GB | 16.1 GB | The community default. Best quality-per-gigabyte for most people. |
| Q4_0 | 14.1 GB | 0.3 GB | 15.2 GB | Older-style 4-bit, measurably worse than Q4_K_M at a similar size. Avoid unless required. |
| IQ4_XS | 13.1 GB | 0.3 GB | 14.2 GB | Newer 4-bit, smaller than Q4_K_M with comparable quality. Needs a recent llama.cpp. |
| Q3_K_M | 12.1 GB | 0.3 GB | 13.2 GB | Noticeable degradation. Use to fit a larger model that would otherwise not run. |
| Q2_K | 10.4 GB | 0.3 GB | 11.4 GB | Heavy degradation. Almost always better to run a smaller model at Q4_K_M instead. |
| IQ3_XS | 10.2 GB | 0.3 GB | 11.3 GB | Aggressive. Usually better than a smaller model at Q4, but test before trusting it. |
Cards that can run it
Best quantisation each card fits at 16k context, with live rental price where we track one.
| GPU | VRAM | Best fit | Uses | ~tok/s | Rent from |
|---|---|---|---|---|---|
| Apple M3 Ultra | 512 GB | F16 | 51 GB | — | — |
| Apple M2 Ultra | 192 GB | F16 | 51 GB | — | — |
| Apple M4 Max (16-core CPU / 40-core GPU) | 128 GB | F16 | 51 GB | — | — |
| AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB) | 128 GB | F16 | 51 GB | — | — |
| NVIDIA DGX Spark (GB10 Grace Blackwell) | 128 GB | F16 | 51 GB | — | — |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | F16 | 51 GB | — | $1.690/hrRunpod |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | F16 | 51 GB | — | — |
| Apple M2 Max | 96 GB | F16 | 51 GB | — | — |
| Apple M4 Pro | 64 GB | F16 | 51 GB | — | — |
| NVIDIA RTX 6000 Ada Generation | 48 GB | Q8_0 | 27 GB | — | $0.539/hrVast.ai |
| AMD Radeon PRO W7900 | 48 GB | Q8_0 | 27 GB | — | — |
| NVIDIA RTX A6000 | 48 GB | Q8_0 | 27 GB | — | $0.330/hrRunpod |
| Apple M4 Max (14-core CPU / 32-core GPU) | 36 GB | Q8_0 | 27 GB | — | — |
| NVIDIA GeForce RTX 5090 | 32 GB | Q8_0 | 27 GB | — | $0.349/hrVast.ai |
| NVIDIA RTX 5000 Ada Generation | 32 GB | Q8_0 | 27 GB | — | $0.336/hrVast.ai |
| NVIDIA GeForce RTX 4090 | 24 GB | Q6_K | 21 GB | — | $0.136/hrVast.ai |
| AMD Radeon RX 7900 XTX | 24 GB | Q6_K | 21 GB | — | — |
| NVIDIA GeForce RTX 3090 Ti | 24 GB | Q6_K | 21 GB | — | $0.121/hrVast.ai |
| NVIDIA RTX A5000 | 24 GB | Q6_K | 21 GB | — | $0.160/hrRunpod |
| NVIDIA GeForce RTX 3090 | 24 GB | Q6_K | 21 GB | — | — |
| AMD Radeon RX 7900 XT | 20 GB | Q5_K_Mtight | 19 GB | — | — |
| NVIDIA GeForce RTX 5060 Ti 16GB | 16 GB | Q4_0tight | 15 GB | — | $0.074/hrVast.ai |
| AMD Radeon RX 9070 XT | 16 GB | Q4_0tight | 15 GB | — | — |
| AMD Radeon RX 9070 | 16 GB | Q4_0tight | 15 GB | — | — |
| NVIDIA GeForce RTX 4060 Ti 16GB | 16 GB | Q4_0tight | 15 GB | — | $0.068/hrVast.ai |
| NVIDIA GeForce RTX 4070 Ti SUPER | 16 GB | Q4_0tight | 15 GB | — | $0.164/hrVast.ai |
| NVIDIA GeForce RTX 4080 | 16 GB | Q4_0tight | 15 GB | — | $0.295/hrVast.ai |
| NVIDIA GeForce RTX 4080 SUPER | 16 GB | Q4_0tight | 15 GB | — | $0.187/hrVast.ai |
| NVIDIA RTX A4000 | 16 GB | Q4_0tight | 15 GB | — | $0.072/hrVast.ai |
| AMD Radeon RX 9060 XT 16GB | 16 GB | Q4_0tight | 15 GB | — | — |
| NVIDIA GeForce RTX 5080 | 16 GB | Q4_0tight | 15 GB | — | $0.201/hrVast.ai |
| NVIDIA GeForce RTX 5070 Ti | 16 GB | Q4_0tight | 15 GB | — | $0.122/hrVast.ai |
| NVIDIA GeForce RTX 4070 Ti | 12 GB | Q2_Ktight | 11 GB | — | — |
| NVIDIA GeForce RTX 4070 SUPER | 12 GB | Q2_Ktight | 11 GB | — | $0.096/hrVast.ai |
| NVIDIA GeForce RTX 4070 | 12 GB | Q2_Ktight | 11 GB | — | $0.165/hrVast.ai |
| NVIDIA GeForce RTX 3080 Ti | 12 GB | Q2_Ktight | 11 GB | — | $0.142/hrVast.ai |
| NVIDIA GeForce RTX 3060 12GB | 12 GB | Q2_Ktight | 11 GB | — | $0.053/hrVast.ai |
| NVIDIA GeForce RTX 5070 | 12 GB | Q2_Ktight | 11 GB | — | — |
| NVIDIA GeForce RTX 3080 12GB | 12 GB | Q2_Ktight | 11 GB | — | $0.081/hrVast.ai |
How these numbers are produced
- Model shape comes from the model's own config.json — 5 layers, 16 attention heads, 2 KV heads.
- Weight size is parameters × effective bits-per-weight. Those constants are checked against real published quantised file sizes — median error 0.7% across 40 measurements.
- KV cache is 2 × layers × KV-heads × head-dim × context × 2 bytes. Using KV-heads rather than attention heads is what makes this correct for grouped-query attention; treating a GQA model as multi-head overstates the cache by up to 8×.
- Tokens/sec is an estimate, not a benchmark. Generation is memory-bandwidth-bound, so this is bandwidth ÷ weight-bytes derated to 75%. Real throughput depends on your runtime, batch size and kernels. We don't run our own hardware tests — see our editorial policy.
- This is a mixture-of-experts model. Generation reads only the experts active for each token (8 of 128), so throughput is scored on ~3.8B active parameters, not the full 27B. Memory is the opposite: every expert must still be resident in VRAM, so the figures above use the full weight set.Hybrid attention: 30 layers, 5 full (global: 2 KV heads × 512 head_dim) + 25 sliding-window (1,024 tokens: 8 KV heads × 256, a fixed ~0.2 GB at fp16). `layers`/kv_heads/head_dim describe the 5 global layers so long-context KV is right. Active count = vendor-published 3.8B. Multimodal weights included in params. Accessed 25 Sep 2026.
- Some memory-bandwidth figures are not yet independently verified.Where that's the case we show no tokens/sec at all rather than a number we can't stand behind.
- Rental prices are pulled hourly from provider APIs (last updated Thu, 24 Sep 2026 21:20:58 GMT). Spot/interruptible pricing can change or vanish without notice.
Try it interactively
Drag the context slider and watch the KV cache fill the card
Worth reading before you buy
GLM-4.5-Air on local hardware
The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.
Cheapest way to run GLM in the cloud
Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.
Run Mistral Small 24B locally
VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.
Cheapest RTX 4090 rental
Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.