ALITEQ.

The Best GPU for Running Qwen3 Locally in 2026 (by Model Size)

Qwen3 isn't one model — it's a family from 4B to 235B, and the right GPU depends entirely on which you run. My VRAM-first map of every variant, the MoE trap that trips people up, and the used card I'd hand most people.

Lena FischerUpdated 54m ago8 min readWeb story
A modern graphics card glowing inside a dark PC build, for running Qwen3 locally
Share

People ask me which GPU to buy for Qwen3 as if it's one model, and it isn't — that's the whole trick to answering this well. Qwen3 is a family that runs from a 0.6B you could load on a phone to a 235B mixture-of-experts that won't fit in any single consumer card. So the real question isn't 'best GPU for Qwen3,' it's 'which Qwen3 do you actually want to run,' and once you answer that, the card picks itself. I run these models daily, and my rule is the same one I give for everything local: size the VRAM to the variant, then buy the cheapest card that clears it. Here's the whole family mapped to memory, and the one card I'd hand most people.

~6GB

Qwen3-8B

Q4 + KV cache

~10GB

Qwen3-14B

Q4 + KV cache

~20GB

Qwen3-32B

Q4 — a 24GB card

~130GB+

Qwen3-235B

full weights — not one GPU

A modern dual-fan graphics card glowing inside a dark PC build
There's no single 'Qwen3 GPU' — there's the right card for the size you run. Illustration generated with AI. · Generated with Higgsfield

Map the Qwen3 you want to the VRAM it needs

Qwen3's open lineup is six dense models (0.6B, 1.7B, 4B, 8B, 14B, 32B) plus two mixture-of-experts models (30B-A3B and 235B-A22B). For local use, the numbers that matter are the memory each needs at Q4_K_M — the quantization almost everyone runs — plus roughly 1–2GB for the KV cache and runtime. Based on the public VRAM calculators and deployment guides, the 4B lands around 2.5GB, the 8B around 5–6GB, the 14B around 8–10GB, and the 32B dense around 18–20GB. The MoE models are the ones people get wrong, so they get their own section below. In my view the 8B and 14B are the everyday sweet spot — genuinely useful, and they fit a mid-range card with room for context.

Qwen3 variant → VRAM at Q4 → the card I'd run it on

Qwen3-4B

VRAM at Q4 (+KV cache)
~3–4GB
Card I'd use
Anything, even a laptop / Arc

Qwen3-8B

VRAM at Q4 (+KV cache)
~6GB
Card I'd use
8–12GB: Arc B580, used 3060 12GB

Qwen3-14B

VRAM at Q4 (+KV cache)
~10GB
Card I'd use
12–16GB: RX 9070, RTX 5070 Ti

Qwen3-32B (dense)

VRAM at Q4 (+KV cache)
~20GB
Card I'd use
24GB: used RTX 3090 / 4090

Qwen3-30B-A3B (MoE)

VRAM at Q4 (+KV cache)
~18–20GB
Card I'd use
24GB card — full weights load

Qwen3-235B-A22B (MoE)

VRAM at Q4 (+KV cache)
~130GB+
Card I'd use
Cloud / unified memory / multi-GPU

The MoE trap: plan for the total size, not the active experts

Here's the mistake I see constantly. Qwen3-235B-A22B activates only 22B parameters per token, so people assume they can run it like a 22B model. You can't. All 235B parameters have to sit in memory — the router picks which experts fire per token, but they all have to be loaded and ready. At Q4 that's roughly 130GB-plus, which is firmly datacenter, unified-memory, or multi-GPU territory. The same logic applies to the friendlier 30B-A3B: only 3B active, but you still need room for all 30B, so treat it as a 24GB-card model. Plan hardware on the total parameter count, every time.

The value pick almost nobody regrets: a used RTX 3090

If I could hand you one card for Qwen3 and walk away, it'd be a used RTX 3090. It has 24GB of VRAM — the same as a used 4090 — so it runs exactly the same models at the same quality, including Qwen3-32B and the 30B-A3B MoE. The difference is only speed, and the price gap is enormous: reviewers peg a used 3090 at roughly $700–1,000 versus $2,000+ for a used 4090. For inference you're getting most of the capability for well under half the money, and 24GB is the single most useful amount of memory for the local-AI hobbyist right now. Buy from a seller with returns, and if you want the fuller reasoning I laid it out in the used-GPU value picks.

24GB of VRAM, two ways — used-market price (Sep 2026)

Used RTX 3090 (24GB)~$700–1,000

same models, same quality

Used RTX 4090 (24GB)~$2,000+

faster, not more capable

Verdict

What I'd buy for Qwen3

Running 8B–14B (most people): a 16GB card — RX 9070 on value, RTX 5070 Ti if you want CUDA's smoother software. Running 32B or the 30B-A3B MoE: get 24GB, and a used RTX 3090 is the value king — same capability as a 4090 at under half the price. Wanting 235B: that's not a graphics-card question — rent it or use a unified-memory box. Whatever you pick, size the VRAM to the exact variant first; the rest is just how fast it runs.

Best for: Anyone choosing a GPU specifically to run Qwen3 models locally

Common questions

Can I run Qwen3-235B-A22B locally?
Only on serious hardware. Despite activating just 22B parameters per token, all 235B must be in memory — roughly 130GB+ at Q4. That means a unified-memory machine (128GB+ shared), a multi-GPU rig, or a rented datacenter GPU. No single consumer graphics card can do it.
What GPU runs Qwen3-32B locally?
A 24GB card. Qwen3-32B needs ~20GB at Q4 plus KV cache, which is comfortable on 24GB and too tight on 16GB for real use. A used RTX 3090 (24GB, ~$700–1,000) is the value pick; a used or new 4090 is the faster option at the same capability.
Is 16GB enough for Qwen3?
Yes for the everyday models — Qwen3-8B and 14B run comfortably on 16GB at Q4 with context to spare, and you can squeeze a heavily-quantized 32B if you accept the quality hit. For full-quality 32B, step up to 24GB.
Used RTX 3090 or a new mid-range card for Qwen3?
For local AI, the 3090's 24GB usually wins — VRAM is the ceiling on what you can run, and 24GB clears everything up to 32B-class. A new 16GB card is a fine buy if you only run 8B–14B and want a warranty and lower power draw.

This is one model's version of a question I answer for every model — the full framework, and the cards I'd buy at each budget, live in my best-GPU-for-local-AI pillar. Size your specific target in the cost-to-run tool, see the wider field in the best 16GB local-LLM GPUs, and if one card won't cut it, read whether two cheaper cards beat one big one.

Found this useful? Share it

Share
Lena Fischer

AI & Local Compute Editor

Lena Fischer

Lena runs more GPUs at home than she'll admit to and has quantized more models than she's finished reading about. She writes about running AI on your own hardware — what actually fits, what's genuinely fast, and what the polished cloud demos quietly leave out.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading