What it costs to run Mixtral 8x7B (MoE)
46.7B paramsMoE · 8 experts, 2 active32k max contextapache-2.0
Memory needed, by quantisation
At 16k context. Weights + KV cache + ~0.8 GB overhead.
| Quant | Weights | KV cache | Total | Quality |
|---|---|---|---|---|
| F16 | 87.0 GB | 2.0 GB | 89.8 GB | Full precision. Reference quality, twice the size of Q8 for no practical gain in most chat use. |
| Q8_0 | 46.2 GB | 2.0 GB | 49.0 GB | Effectively lossless. Use when VRAM is not the constraint. |
| Q6_K | 35.7 GB | 2.0 GB | 38.4 GB | Very close to Q8 at meaningfully less memory. A safe high-quality pick. |
| Q5_K_M | 30.9 GB | 2.0 GB | 33.7 GB | Small, hard-to-notice quality loss. Good balance. |
| Q5_0 | 30.1 GB | 2.0 GB | 32.9 GB | Older-style 5-bit. Q5_K_M is usually the better pick at the same size. |
| Q4_K_M | 26.4 GB | 2.0 GB | 29.2 GB | The community default. Best quality-per-gigabyte for most people. |
| Q4_0 | 24.7 GB | 2.0 GB | 27.5 GB | Older-style 4-bit, measurably worse than Q4_K_M at a similar size. Avoid unless required. |
| IQ4_XS | 23.1 GB | 2.0 GB | 25.9 GB | Newer 4-bit, smaller than Q4_K_M with comparable quality. Needs a recent llama.cpp. |
| Q3_K_M | 21.3 GB | 2.0 GB | 24.0 GB | Noticeable degradation. Use to fit a larger model that would otherwise not run. |
| Q2_K | 18.2 GB | 2.0 GB | 21.0 GB | Heavy degradation. Almost always better to run a smaller model at Q4_K_M instead. |
| IQ3_XS | 17.9 GB | 2.0 GB | 20.7 GB | Aggressive. Usually better than a smaller model at Q4, but test before trusting it. |
Cards that can run it
Best quantisation each card fits at 16k context, with live rental price where we track one.
| GPU | VRAM | Best fit | Uses | ~tok/s | Rent from |
|---|---|---|---|---|---|
| Apple M3 Ultra | 512 GB | F16 | 90 GB | ~24 | — |
| Apple M2 Ultra | 192 GB | F16 | 90 GB | ~24 | — |
| Apple M4 Max (16-core CPU / 40-core GPU) | 128 GB | F16 | 90 GB | ~16 | — |
| AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB) | 128 GB | F16 | 90 GB | ~7 | — |
| NVIDIA DGX Spark (GB10 Grace Blackwell) | 128 GB | F16 | 90 GB | ~8 | — |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | F16tight | 90 GB | ~52 | $1.690/hr |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | F16tight | 90 GB | ~52 | — |
| Apple M2 Max | 96 GB | F16tight | 90 GB | ~12 | — |
| Apple M4 Pro | 64 GB | Q8_0 | 49 GB | ~15 | — |
| NVIDIA RTX 6000 Ada Generation | 48 GB | Q6_K | 38 GB | ~68 | $0.428/hr |
| NVIDIA RTX A6000 | 48 GB | Q6_K | 38 GB | ~55 | $0.330/hr |
| AMD Radeon PRO W7900 | 48 GB | Q6_K | 38 GB | ~61 | — |
| Apple M4 Max (14-core CPU / 32-core GPU) | 36 GB | Q5_K_Mtight | 34 GB | ~34 | — |
| NVIDIA RTX 5000 Ada Generation | 32 GB | Q4_K_Mtight | 29 GB | ~55 | $0.490/hr |
| NVIDIA GeForce RTX 5090 | 32 GB | Q4_K_Mtight | 29 GB | ~172 | $0.058/hr |
| NVIDIA RTX A5000 | 24 GB | Q2_K | 21 GB | ~107 | $0.160/hr |
| NVIDIA GeForce RTX 4090 | 24 GB | Q2_K | 21 GB | ~140 | $0.136/hr |
| NVIDIA GeForce RTX 3090 Ti | 24 GB | Q2_K | 21 GB | ~140 | $0.116/hr |
| AMD Radeon RX 7900 XTX | 24 GB | Q2_K | 21 GB | ~134 | — |
| NVIDIA GeForce RTX 3090 | 24 GB | Q2_K | 21 GB | ~130 | — |
How these numbers are produced
- Model shape comes from the model's own config.json — 32 layers, 32 attention heads, 8 KV heads.
- Weight size is parameters × effective bits-per-weight. Those constants are checked against real published quantised file sizes — median error 0.7% across 40 measurements.
- KV cache is 2 × layers × KV-heads × head-dim × context × 2 bytes. Using KV-heads rather than attention heads is what makes this correct for grouped-query attention; treating a GQA model as multi-head overstates the cache by up to 8×.
- Tokens/sec is an estimate, not a benchmark. Generation is memory-bandwidth-bound, so this is bandwidth ÷ weight-bytes derated to 75%. Real throughput depends on your runtime, batch size and kernels. We don't run our own hardware tests — see our editorial policy.
- This is a mixture-of-experts model. Generation reads only the experts active for each token (2 of 8), so throughput is scored on ~12.9B active parameters, not the full 47B. Memory is the opposite: every expert must still be resident in VRAM, so the figures above use the full weight set.Derived from architecture; matches the vendor-published active count within 1%.
- Rental prices are pulled hourly from provider APIs (last updated Tue, 21 Jul 2026 15:20:05 GMT). Spot/interruptible pricing can change or vanish without notice.