Mistral Small 24B fits a single 24GB card at Q4 (~14GB), running ~85 tok/s on a 4090. VRAM by quant, which GPU, and the cheapest way to run it — own or rent.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
Mistral Small 24B is one of the best models you can run on a single consumer GPU: a 24-billion-parameter dense model that punches well above its size, and at 4-bit quantisation it fits a 24GB card with room to spare. Here's exactly what it needs — VRAM by quant, which GPU, what speed to expect — and the cheapest way to run it, whether you own a card or rent one. Part of our local-AI hardware guides.
VRAM by quantisation
Mistral Small 24B — VRAM by quant
Q4_K_M
VRAM (approx)
~14–15GB
Fits
16GB card (tight), 24GB (comfortable)
Trade-off
Best value; tiny quality cost
Q6_K
VRAM (approx)
~19–20GB
Fits
24GB card
Trade-off
Higher quality
Q8_0
VRAM (approx)
~25GB
Fits
24GB (tight) / 32GB
Trade-off
Near-lossless
BF16 (full)
VRAM (approx)
~55GB
Fits
A100 80GB / 2×24GB
Trade-off
No quality loss, rarely needed
VRAM (approx)
Fits
Trade-off
Q4_K_M
~14–15GB
16GB card (tight), 24GB (comfortable)
Best value; tiny quality cost
Q6_K
~19–20GB
24GB card
Higher quality
Q8_0
~25GB
24GB (tight) / 32GB
Near-lossless
BF16 (full)
~55GB
A100 80GB / 2×24GB
No quality loss, rarely needed
No 24GB card? Rent one from pennies an hourReferral link
For the best value, a 24GB card — a used RTX 3090 or an RTX 4090 — runs Q4_K_M with about 3.6GB of headroom for context, at roughly 85 tokens/second on a 4090 (published benchmark). A 16GB card (like a 5060 Ti 16GB) fits Q4 but with little room for long context. For Q8 or a big context window you'll want 32GB+ or a data-centre card. See the model-specific pick on our best GPU for Mistral Small 24B page and match any model to a card with the cost-to-run tool.
Or run it in the cloud
If you don't own a 24GB card, renting one is the cheapest way to try Mistral Small 24B at full speed: a 24GB GPU is about $0.14–0.34/hour as of 24 Sep 2026, so an evening of use costs a dollar or two and you buy nothing. It's also the honest way to test whether the model is worth a hardware purchase before you commit — rent first, then decide. Our how to run Mistral locally guide covers the software side.
Referral link
Run Mistral Small 24B on a rented 24GB card
A 24GB card runs it at Q4 comfortably — from about $0.14/hr on Vast.ai spot (24 Sep 2026). Pay pennies for an evening; buy nothing.
Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.
Frequently asked
How much VRAM does Mistral Small 24B need?
About 14–15GB at Q4_K_M (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card with headroom for context. Q8 needs ~25GB and full BF16 about 55GB. For most people, Q4_K_M on a 24GB card is the sweet spot.
Can an RTX 4090 run Mistral Small 24B?
Yes, easily — a 24GB RTX 4090 runs Q4_K_M with roughly 3.6GB of headroom for context, at about 85 tokens/second (published benchmark). A used RTX 3090 (also 24GB) runs it too, a bit slower. Both are ideal single-card homes for this model.
What's the cheapest way to run Mistral Small 24B if I don't have a good GPU?
Rent a 24GB card in the cloud — from about $0.14/hour on Vast.ai spot or $0.34/hour on Runpod as of 24 September 2026. You pay only while it runs, so trying the model costs a dollar or two, and you can decide whether it's worth buying hardware before you spend.
Referral link
Managed pods + serverless
Prefer a stable, managed box? Runpod runs a 24GB card at about $0.34/hr on-demand (24 Sep 2026).