aliteq.

run Mistral Small 24B locally in 2026: VRAM, which GPU, and how fast

Mistral Small 24B fits a single 24GB card at Q4 (~14GB). VRAM by quant, which GPU, realistic speeds by runtime, and the cheapest way to run it — own or rent.

TensorUpdated Sep 247 min readWeb story
Flat illustration of a person running a local AI chat assistant on a laptop with a single graphics card, on a teal background

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

Mistral Small 24B is one of the best models you can run on a single consumer GPU: a 24-billion-parameter dense model that punches well above its size, and at 4-bit quantisation it fits a 24GB card with room to spare. Here's exactly what it needs — VRAM by quant, which GPU, what speed to expect — and the cheapest way to run it, whether you own a card or rent one. Part of our local-AI hardware guides.

VRAM by quantisation

Mistral Small 24B — VRAM by quant

Q4_K_M

VRAM (approx)
~14–15GB
Fits
16GB card (tight), 24GB (comfortable)
Trade-off
Best value; tiny quality cost

Q6_K

VRAM (approx)
~19–20GB
Fits
24GB card
Trade-off
Higher quality

Q8_0

VRAM (approx)
~25GB
Fits
24GB (tight) / 32GB
Trade-off
Near-lossless

BF16 (full)

VRAM (approx)
~55GB
Fits
A100 80GB / 2×24GB
Trade-off
No quality loss, rarely needed
Vast.ai

No 24GB card? Rent one from pennies an hourReferral link

Rent a 24GB card

Which GPU to run it on

For the best value, a 24GB card — a used RTX 3090 or an RTX 4090 — runs Q4_K_M with about 3.6GB of headroom for context, at about 38 tokens/second with llama.cpp, or roughly 85 with vLLM and AWQ INT4 (Markaicode, citing community benchmarks). A 16GB card (like a 5060 Ti 16GB) fits Q4 but with little room for long context. For Q8 or a big context window you'll want 32GB+ or a data-centre card. See the model-specific pick on our best GPU for Mistral Small 24B page and match any model to a card with the cost-to-run tool.

Or run it in the cloud

If you don't own a 24GB card, renting one is the cheapest way to try Mistral Small 24B at full speed: a 24GB GPU is about $0.14–0.34/hour as of 24 Sep 2026, so an evening of use costs a dollar or two and you buy nothing. It's also the honest way to test whether the model is worth a hardware purchase before you commit — rent first, then decide. Our how to run Mistral locally guide covers the software side.

Vast.aiReferral link

Run Mistral Small 24B on a rented 24GB card

A 24GB card runs it at Q4 comfortably — from about $0.14/hr on Vast.ai spot (24 Sep 2026). Pay pennies for an evening; buy nothing.

Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.

Frequently asked

How much VRAM does Mistral Small 24B need?
About 14–15GB at Q4_K_M (a ~14GB GGUF), so it fits a 16GB card tightly and runs comfortably on a 24GB card with headroom for context. Q8 needs ~25GB and full BF16 about 55GB. For most people, Q4_K_M on a 24GB card is the sweet spot.
Can an RTX 4090 run Mistral Small 24B?
Yes, easily — a 24GB RTX 4090 runs Q4_K_M with roughly 3.6GB of headroom for context, at about 38 tokens/second with llama.cpp at Q4_K_M, or about 85 with vLLM and AWQ INT4, per community benchmarks compiled by Markaicode. A used RTX 3090 (also 24GB) runs it too, a bit slower. Both are ideal single-card homes for this model.
What's the cheapest way to run Mistral Small 24B if I don't have a good GPU?
Rent a 24GB card in the cloud — from about $0.14/hour on Vast.ai spot or $0.34/hour on Runpod as of 24 September 2026. You pay only while it runs, so trying the model costs a dollar or two, and you can decide whether it's worth buying hardware before you spend.
RunpodReferral link

Managed pods + serverless

Prefer a stable, managed box? Runpod runs a 24GB card at about $0.34/hr on-demand (24 Sep 2026).

Mistral Small 24B is the rare model that's genuinely great on one consumer card: Q4 on a 24GB GPU is all most people need. Own a card or rent one cheaply; see the best GPU for Mistral Small 24B and how to run Mistral locally.

Found this useful? Share it

Share
Tensor

Local AI & Automation Editor

Tensor

I'm US-based, I run more models at home than I'll admit to, and I've quantized more than I've finished reading about. I write about running AI on your own hardware and, lately, about what it costs a company to do the same — tokens per day, GPUs per month, and the GDPR questions nobody's sales deck answers.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading