Qwen3-32B at Q4 needs about 20GB of VRAM, which rules out every 16GB card and makes this a simple question: what's the cheapest 24GB GPU that runs it well? The answer is used.
Ask 'what's the best GPU for Qwen3-32B' and the honest answer is anticlimactic: whichever 24GB card is cheapest, because the model's requirement makes speed the tiebreaker, not the decider. Qwen3-32B at Q4_K_M occupies roughly 20GB before context — comfortably inside 24GB, impossible in 16GB. So the whole field narrows to 24GB cards, and among those, a used RTX 3090 runs the model at ~30 tokens/sec for about $1,000. That's the pick for most people, and here's the reasoning in full.
The 20GB wall
Everything about this decision follows from one number. A 32-billion-parameter model at Q4_K_M quantization is roughly 20GB of weights, and you need headroom on top for the context window's KV cache. That fits a 24GB card with a little room to spare and overflows a 16GB card completely. There's no clever setting that changes the physics — the model either fits your VRAM or it spills to system RAM and crawls. So the first and only hard filter is 24GB, and then it's a value question among the cards that clear it. Check your exact target in the VRAM calculator if you plan to run long context.
Does it fit? — Qwen3-32B (Q4) needs ≈20GB
Qwen3-32B (Q4) needs ≈20 GB
RTX 5060 Ti (16GB)16 GBover 4 GB
RTX 3090 (24GB)24 GBfits
RTX 4090 (24GB)24 GBfits
RTX 5090 (32GB)32 GBfits
16GB cards fall short; every 24GB card clears it. That's the whole shortlist.
Which 24GB card
The value pick vs the fast pick
Used RTX 3090
24GB · ~$1,000
vs
RTX 4090
24GB · ~$2,200
Yes, ~30 t/s
Runs Qwen3-32B
Yes, faster
24 GB
VRAM
24 GB
~$1,000
Price
~$2,200
Good
Speed
Better
Best
Value for this model
Overkill
3090 wins 2wins 1 4090
For Qwen3-32B specifically, a used 3090's 24GB does everything a pricier card does — just a bit slower. · Unsplash
Verdict
Used RTX 3090 for value, 4090/5090 for speed
Qwen3-32B is a capacity problem with a value answer. Any 24GB card runs it; a used RTX 3090 at ~$1,000 does so at a genuinely usable ~30 tokens/sec, making it the best-value choice for this model. Step up to a 4090 or 5090 only if you want more speed or plan to run bigger models later. Skip 16GB cards entirely for Qwen3-32B — they can't hold it.
Best for: Used 3090: best value, 24GB, ~30 t/s. 4090/5090: more speed, more money. 16GB cards: pick a 14B model instead.
Quick answers
Can I run Qwen3-32B on a 16GB GPU?
Not properly. At Q4 the model needs about 20GB, so a 16GB card can't hold it in VRAM. You could offload part of it to system RAM, but that drops you to a few tokens per second — too slow for real use. On 16GB you're better off running a 14B-class model that fits comfortably. For Qwen3-32B specifically, 24GB is the minimum that makes sense.
How fast is Qwen3-32B on a 3090?
Around 30 tokens per second at Q4 with 16k context, per hardware-corner's llama.cpp benchmarks — comfortably usable for chat and coding. A 4090 is faster and a 5090 faster still, but 30 t/s on a $1,000 used card is the value sweet spot. Speed scales with the card, but capability (does it run at all) is identical across every 24GB option.
Is the 5090's 32GB worth it for Qwen3-32B?
Only if you value the extra speed or want headroom for larger models and long context. For Qwen3-32B alone, the 32GB isn't necessary — 24GB holds it fine. The 5090's advantage is speed and future-proofing, not the ability to run this particular model. If Qwen3-32B is your ceiling, a cheaper 24GB card is the smarter spend.
Bottom line: for Qwen3-32B, buy the cheapest 24GB card you're comfortable with, and a used 3090 is that card for most people. Size it in the VRAM calculator, and if you want the general 24GB model landscape, our best local LLM on 24GB guide covers what else fits alongside it.