run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast
Mistral Small 24B fits a single 24GB card at Q4 (~14GB). VRAM by quant, which GPU, realistic speeds by runtime, and the cheapest way to run it — own…
Aliteq
Tensor · Local AI & Automation Editor
The short answer
Mistral Small 24B (v3.2) is a 24B dense model. At Q4_K_M it needs about 14–15GB of VRAM (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card like an RTX 3090 or 4090…
Q4_K_M: ~14–15GB — fits a 16GB card (tight), comfortable on 24GB.
Q8_0: ~25GB — needs a 24GB card at minimum, better on 32GB+.
Full BF16: ~55GB — an A100 80GB or two 24GB cards.
Speed on a 4090: ~38 tok/s (llama.cpp, Q4_K_M) or ~85 tok/s (vLLM, AWQ INT4), per Markaicode's compilation of community benchmarks.
No capable GPU? Rent a 24GB card from ~$0.14/hr and run it in the cloud.
Aliteq
Read the full story
run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast