What it costs to run Qwen3 14B
Trying to decide what to buy? The best GPU for Qwen3 14B →
Cheapest way to rent it right now
$0.05/hr
NVIDIA GeForce RTX 3060 12GB · Q4_K_M · spot · ≈ $13/mo at 8h/day
Live price captured Sat, 10 Oct 2026 11:21:09 GMT. Referral link — we may earn a commission at no cost to you; it never changes which card is cheapest.
Memory needed, by quantisation
At 16k context. Weights + KV cache + ~0.8 GB overhead.
| Quant | Weights | KV cache | Total | Quality |
|---|---|---|---|---|
| F16 | 27.5 GB | 2.5 GB | 30.8 GB | Full precision. Reference quality, twice the size of Q8 for no practical gain in most chat use. |
| Q8_0 | 14.6 GB | 2.5 GB | 17.9 GB | Effectively lossless. Use when VRAM is not the constraint. |
| Q6_K | 11.3 GB | 2.5 GB | 14.6 GB | Very close to Q8 at meaningfully less memory. A safe high-quality pick. |
| Q5_K_M | 9.8 GB | 2.5 GB | 13.0 GB | Small, hard-to-notice quality loss. Good balance. |
| Q5_0 | 9.5 GB | 2.5 GB | 12.8 GB | Older-style 5-bit. Q5_K_M is usually the better pick at the same size. |
| Q4_K_M | 8.3 GB | 2.5 GB | 11.6 GB | The community default. Best quality-per-gigabyte for most people. |
| Q4_0 | 7.8 GB | 2.5 GB | 11.1 GB | Older-style 4-bit, measurably worse than Q4_K_M at a similar size. Avoid unless required. |
| IQ4_XS | 7.3 GB | 2.5 GB | 10.6 GB | Newer 4-bit, smaller than Q4_K_M with comparable quality. Needs a recent llama.cpp. |
| Q3_K_M | 6.7 GB | 2.5 GB | 10.0 GB | Noticeable degradation. Use to fit a larger model that would otherwise not run. |
| Q2_K | 5.8 GB | 2.5 GB | 9.0 GB | Heavy degradation. Almost always better to run a smaller model at Q4_K_M instead. |
| IQ3_XS | 5.7 GB | 2.5 GB | 9.0 GB | Aggressive. Usually better than a smaller model at Q4, but test before trusting it. |
Cards that can run it
Best quantisation each card fits at 16k context, with live rental price where we track one.
| GPU | VRAM | Best fit | Uses | ~tok/s | Rent from |
|---|---|---|---|---|---|
| Apple M3 Ultra | 512 GB | F16 | 31 GB | — | — |
| Apple M2 Ultra | 192 GB | F16 | 31 GB | — | — |
| Apple M4 Max (16-core CPU / 40-core GPU) | 128 GB | F16 | 31 GB | — | — |
| AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB) | 128 GB | F16 | 31 GB | — | — |
| NVIDIA DGX Spark (GB10 Grace Blackwell) | 128 GB | F16 | 31 GB | — | — |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | F16 | 31 GB | — | $1.690/hrRunpod |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | F16 | 31 GB | — | — |
| Apple M2 Max | 96 GB | F16 | 31 GB | — | — |
| Apple M4 Pro | 64 GB | F16 | 31 GB | — | — |
| NVIDIA RTX 6000 Ada Generation | 48 GB | F16 | 31 GB | — | $0.494/hrVast.ai |
| AMD Radeon PRO W7900 | 48 GB | F16 | 31 GB | — | — |
| NVIDIA RTX A6000 | 48 GB | F16 | 31 GB | — | $0.281/hrVast.ai |
| Apple M4 Max (14-core CPU / 32-core GPU) | 36 GB | F16 | 31 GB | — | — |
| NVIDIA GeForce RTX 5090 | 32 GB | F16tight | 31 GB | — | $0.374/hrVast.ai |
| NVIDIA RTX 5000 Ada Generation | 32 GB | F16tight | 31 GB | — | $0.336/hrVast.ai |
| NVIDIA GeForce RTX 4090 | 24 GB | Q8_0 | 18 GB | — | $0.340/hrRunpod |
| AMD Radeon RX 7900 XTX | 24 GB | Q8_0 | 18 GB | — | — |
| NVIDIA GeForce RTX 3090 Ti | 24 GB | Q8_0 | 18 GB | — | $0.120/hrVast.ai |
| NVIDIA RTX A5000 | 24 GB | Q8_0 | 18 GB | — | $0.160/hrRunpod |
| NVIDIA GeForce RTX 3090 | 24 GB | Q8_0 | 18 GB | — | — |
| AMD Radeon RX 7900 XT | 20 GB | Q8_0 | 18 GB | — | — |
| NVIDIA GeForce RTX 5060 Ti 16GB | 16 GB | Q6_Ktight | 15 GB | — | $0.099/hrVast.ai |
| AMD Radeon RX 9070 XT | 16 GB | Q6_Ktight | 15 GB | — | — |
| AMD Radeon RX 9070 | 16 GB | Q6_Ktight | 15 GB | — | — |
| NVIDIA GeForce RTX 4060 Ti 16GB | 16 GB | Q6_Ktight | 15 GB | — | $0.069/hrVast.ai |
| NVIDIA GeForce RTX 4070 Ti SUPER | 16 GB | Q6_Ktight | 15 GB | — | $0.149/hrVast.ai |
| NVIDIA GeForce RTX 4080 | 16 GB | Q6_Ktight | 15 GB | — | $0.215/hrVast.ai |
| NVIDIA GeForce RTX 4080 SUPER | 16 GB | Q6_Ktight | 15 GB | — | $0.201/hrVast.ai |
| NVIDIA RTX A4000 | 16 GB | Q6_Ktight | 15 GB | — | $0.066/hrVast.ai |
| AMD Radeon RX 9060 XT 16GB | 16 GB | Q6_Ktight | 15 GB | — | — |
| NVIDIA GeForce RTX 5080 | 16 GB | Q6_Ktight | 15 GB | — | $0.267/hrVast.ai |
| NVIDIA GeForce RTX 5070 Ti | 16 GB | Q6_Ktight | 15 GB | — | $0.122/hrVast.ai |
| NVIDIA GeForce RTX 4070 Ti | 12 GB | Q4_K_Mtight | 12 GB | — | — |
| NVIDIA GeForce RTX 4070 SUPER | 12 GB | Q4_K_Mtight | 12 GB | — | $0.094/hrVast.ai |
| NVIDIA GeForce RTX 4070 | 12 GB | Q4_K_Mtight | 12 GB | — | $0.190/hrVast.ai |
| NVIDIA GeForce RTX 3080 Ti | 12 GB | Q4_K_Mtight | 12 GB | — | $0.121/hrVast.ai |
| NVIDIA GeForce RTX 3060 12GB | 12 GB | Q4_K_Mtight | 12 GB | — | $0.053/hrVast.ai |
| NVIDIA GeForce RTX 5070 | 12 GB | Q4_K_Mtight | 12 GB | — | — |
| NVIDIA GeForce RTX 3080 12GB | 12 GB | Q4_K_Mtight | 12 GB | — | $0.080/hrVast.ai |
| NVIDIA GeForce RTX 3080 10GB | 10 GB | Q2_Ktight | 9 GB | — | — |
How these numbers are produced
- Model shape comes from the model's own config.json — 40 layers, 40 attention heads, 8 KV heads.
- Weight size is parameters × effective bits-per-weight. Those constants are checked against real published quantised file sizes — median error 0.7% across 40 measurements.
- KV cache is 2 × layers × KV-heads × head-dim × context × 2 bytes. Using KV-heads rather than attention heads is what makes this correct for grouped-query attention; treating a GQA model as multi-head overstates the cache by up to 8×.
- Tokens/sec is an estimate, not a benchmark. Generation is memory-bandwidth-bound, so this is bandwidth ÷ weight-bytes derated to 75%. Real throughput depends on your runtime, batch size and kernels. We don't run our own hardware tests — see our editorial policy.
- Some memory-bandwidth figures are not yet independently verified.Where that's the case we show no tokens/sec at all rather than a number we can't stand behind.
- Rental prices are pulled hourly from provider APIs (last updated Sat, 10 Oct 2026 11:21:09 GMT). Spot/interruptible pricing can change or vanish without notice.
Try it interactively
Drag the context slider and watch the KV cache fill the card
Worth reading before you buy
Cheapest RTX 4090 rental
Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.
GLM-4.5-Air on local hardware
The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.
Cheapest way to run GLM in the cloud
Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.
Run Mistral Small 24B locally
VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.