the best GPU for running Qwen3-32B locally isn't the fastest one — it's the cheapest with 24GB

Qwen3-32B at Q4 needs about 20GB of VRAM, which rules out every 16GB card and makes this a simple question: what's the cheapest 24GB GPU that runs…

Aliteq
Ravi Malhotra · Hardware Editor

The short answer

Qwen3-32B at Q4 needs ~20GB of VRAM, so it requires a 24GB GPU — no 16GB card can hold it. The best value is a used RTX 3090 (~$1,000, ~30 tok/s measured). An RTX 4090 or 5090 run it faster at…

Needs ~20GB at Q4 — fits 24GB cards only. 16GB cards can't hold it.

Best value: used RTX 3090 — ~$1,000, ~30 tok/s, 24GB.

Faster: RTX 4090/5090 — more speed, more money, same capacity (or 32GB on the 5090).

Also works: Strix Halo or a Mac via unified memory, slower but efficient.

16GB cards need not apply — you'd drop to a 14B model or offload painfully.

Aliteq

Read the full story

the best GPU for running Qwen3-32B locally isn't the fastest one — it's the cheapest with 24GB

Read the full story on Aliteq