ALITEQ.

The best GPU for running Qwen3 30B-A3B (MoE) (2026)

Ranked from live rental prices and computed VRAM fit, not opinion. Every card is judged at Q4_K_M — the quantisation most people actually run — at 8k context, so it's a fair comparison. A card that can't fit Qwen3 30B-A3B (MoE)at that quality isn't listed, because it could only run a crushed version.

Best value

NVIDIA GeForce RTX 3090 Ti

Best throughput per dollar among cards that fit.

Cheapest that runs it

NVIDIA GeForce RTX 3090 Ti

Lowest hourly rental that fits it — $0.066/hr, Q4_K_M.

No compromise

Most headroom and highest throughput.

Every card that runs it, ranked

Best quantisation each card fits at 8k context, with the cheapest live rental and a bandwidth-derived throughput estimate. Sorted by tokens/sec per dollar.

GPUVRAMVRAM used~tok/sCheapest rentaltok/s per $
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition96 GB19 GB20%$1.690
NVIDIA RTX PRO 6000 Blackwell Workstation Edition96 GB19 GB20%
NVIDIA RTX 6000 Ada Generation48 GB19 GB39%$0.578
NVIDIA RTX A600048 GB19 GB39%$0.287
AMD Radeon PRO W790048 GB19 GB39%
NVIDIA RTX 5000 Ada Generation32 GB19 GB59%$0.428
NVIDIA GeForce RTX 509032 GB19 GB59%$0.334
NVIDIA RTX A500024 GB19 GB78%$0.073
NVIDIA GeForce RTX 409024 GB19 GB78%$0.134
NVIDIA GeForce RTX 3090 Ti24 GB19 GB78%$0.066
AMD Radeon RX 7900 XTX24 GB19 GB78%
NVIDIA GeForce RTX 309024 GB19 GB78%

How this ranking is made — and its limit

Fit and throughput are computed from Qwen3 30B-A3B (MoE)'s own configuration with the engine behind our cost-to-run page (bits-per-weight validated to 0.7% median error). Rental prices are the cheapest live figure across the providers we track, last updated Mon, 10 Aug 2026 19:20:27 GMT.

The value column ranks renting, because that's what we can price precisely. If you're buyinga card, the fit and throughput columns are exactly what you need — but we don't publish a purchase-price value ranking, because we don't have verified street prices and won't invent them. For the buy-vs-rent decision itself, see the guides below.

Tokens/sec is a bandwidth-derived estimate, not a benchmark — we don't run our own hardware tests (editorial policy). This is a mixture-of-experts model without a known active-parameter count, so throughput is omitted rather than guessed.

Check it yourself

See exactly what fits, at any context length

Running Qwen3 30B-A3B (MoE): common questions

How much VRAM do you need to run Qwen3 30B-A3B (MoE)?
At Q4_K_M and 8k context, Qwen3 30B-A3B (MoE) needs about 18.8 GB — 17.2 GB of weights plus 0.8 GB of KV cache and ~0.8 GB overhead. So you want a card with at least that much free memory; the KV cache grows if you use longer context.
Can an RTX 4090 run Qwen3 30B-A3B (MoE)?
Yes. At Q4_K_M it uses about 18.8 GB of the card's 24 GB (78%).
Can an RTX 3090 run Qwen3 30B-A3B (MoE)?
Yes. At Q4_K_M it uses about 18.8 GB of the card's 24 GB (78%).
Can an RTX 4060 Ti 16GB run Qwen3 30B-A3B (MoE)?
Not at Q4_K_M — Qwen3 30B-A3B (MoE) needs more than the card's 16 GB at that quality and 8k context. You'd have to drop to a lower quantisation (worse quality), shorten the context, or use a bigger card or two cards.
What's the cheapest way to run Qwen3 30B-A3B (MoE)?
Among cards that fit it at Q4_K_M, the cheapest to rent right now is the NVIDIA GeForce RTX 3090 Ti at about $0.066/hour. Whether renting or buying is cheaper overall depends on how many hours a day you'll actually use it.

Before you buy

Best GPU for other models