run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast

Mistral Small 24B fits a single 24GB card at Q4 (~14GB), running ~85 tok/s on a 4090. VRAM by quant, which GPU, and the cheapest way to run it — own…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

Mistral Small 24B (v3.2) is a 24B dense model. At Q4_K_M it needs about 14–15GB of VRAM (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card like an RTX 3090 or 4090…

Q4_K_M: ~14–15GB — fits a 16GB card (tight), comfortable on 24GB.

Q8_0: ~25GB — needs a 24GB card at minimum, better on 32GB+.

Full BF16: ~55GB — an A100 80GB or two 24GB cards.

Speed: ~85 tokens/second on an RTX 4090 at 4-bit (published benchmark).

No capable GPU? Rent a 24GB card from ~$0.14/hr and run it in the cloud.

Aliteq

Read the full story

run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast

Read the full story on Aliteq