run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast
Mistral Small 24B fits a single 24GB card at Q4 (~14GB), running ~85 tok/s on a 4090. VRAM by quant, which GPU, and the cheapest way to run it — own…
Aliteq
Lena Fischer · AI & Local Compute Editor
The short answer
Mistral Small 24B (v3.2) is a 24B dense model. At Q4_K_M it needs about 14–15GB of VRAM (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card like an RTX 3090 or 4090…
Q4_K_M: ~14–15GB — fits a 16GB card (tight), comfortable on 24GB.
Q8_0: ~25GB — needs a 24GB card at minimum, better on 32GB+.
Full BF16: ~55GB — an A100 80GB or two 24GB cards.
Speed: ~85 tokens/second on an RTX 4090 at 4-bit (published benchmark).
No capable GPU? Rent a 24GB card from ~$0.14/hr and run it in the cloud.
Aliteq
Read the full story
run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast