aliteq.

run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast

Mistral Small 24B fits a single 24GB card at Q4 (~14GB), running ~85 tok/s on a 4090. VRAM by quant, which GPU, and the cheapest way to run it — own or rent.

Lena FischerUpdated 1h ago7 min readWeb story
Flat illustration of a person running a local AI chat assistant on a laptop with a single graphics card, on a teal background

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

Mistral Small 24B is one of the best models you can run on a single consumer GPU: a 24-billion-parameter dense model that punches well above its size, and at 4-bit quantisation it fits a 24GB card with room to spare. Here's exactly what it needs — VRAM by quant, which GPU, what speed to expect — and the cheapest way to run it, whether you own a card or rent one. Part of our local-AI hardware guides.

VRAM by quantisation

Mistral Small 24B — VRAM by quant

Q4_K_M

VRAM (approx)
~14–15GB
Fits
16GB card (tight), 24GB (comfortable)
Trade-off
Best value; tiny quality cost

Q6_K

VRAM (approx)
~19–20GB
Fits
24GB card
Trade-off
Higher quality

Q8_0

VRAM (approx)
~25GB
Fits
24GB (tight) / 32GB
Trade-off
Near-lossless

BF16 (full)

VRAM (approx)
~55GB
Fits
A100 80GB / 2×24GB
Trade-off
No quality loss, rarely needed
Vast.ai

No 24GB card? Rent one from pennies an hourReferral link

Rent a 24GB card

Which GPU to run it on

For the best value, a 24GB card — a used RTX 3090 or an RTX 4090 — runs Q4_K_M with about 3.6GB of headroom for context, at roughly 85 tokens/second on a 4090 (published benchmark). A 16GB card (like a 5060 Ti 16GB) fits Q4 but with little room for long context. For Q8 or a big context window you'll want 32GB+ or a data-centre card. See the model-specific pick on our best GPU for Mistral Small 24B page and match any model to a card with the cost-to-run tool.

Or run it in the cloud

If you don't own a 24GB card, renting one is the cheapest way to try Mistral Small 24B at full speed: a 24GB GPU is about $0.14–0.34/hour as of 24 Sep 2026, so an evening of use costs a dollar or two and you buy nothing. It's also the honest way to test whether the model is worth a hardware purchase before you commit — rent first, then decide. Our how to run Mistral locally guide covers the software side.

Vast.aiReferral link

Run Mistral Small 24B on a rented 24GB card

A 24GB card runs it at Q4 comfortably — from about $0.14/hr on Vast.ai spot (24 Sep 2026). Pay pennies for an evening; buy nothing.

Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.

Frequently asked

How much VRAM does Mistral Small 24B need?
About 14–15GB at Q4_K_M (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card with headroom for context. Q8 needs ~25GB and full BF16 about 55GB. For most people, Q4_K_M on a 24GB card is the sweet spot.
Can an RTX 4090 run Mistral Small 24B?
Yes, easily — a 24GB RTX 4090 runs Q4_K_M with roughly 3.6GB of headroom for context, at about 85 tokens/second (published benchmark). A used RTX 3090 (also 24GB) runs it too, a bit slower. Both are ideal single-card homes for this model.
What's the cheapest way to run Mistral Small 24B if I don't have a good GPU?
Rent a 24GB card in the cloud — from about $0.14/hour on Vast.ai spot or $0.34/hour on Runpod as of 24 September 2026. You pay only while it runs, so trying the model costs a dollar or two, and you can decide whether it's worth buying hardware before you spend.
RunpodReferral link

Managed pods + serverless

Prefer a stable, managed box? Runpod runs a 24GB card at about $0.34/hr on-demand (24 Sep 2026).

Mistral Small 24B is the rare model that's genuinely great on one consumer card: Q4 on a 24GB GPU is all most people need. Own a card or rent one cheaply; see the best GPU for Mistral Small 24B and how to run Mistral locally.

Found this useful? Share it

Share
Lena Fischer

AI & Local Compute Editor

Lena Fischer

Lena runs more GPUs at home than she'll admit to and has quantized more models than she's finished reading about. She writes about running AI on your own hardware — what actually fits, what's genuinely fast, and what the polished cloud demos quietly leave out.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading