aliteq.

The best GPU for running Qwen3 235B-A22B (MoE) (2026)

Ranked from live rental prices and computed VRAM fit, not opinion. Every card is judged at Q4_K_M — the quantisation most people actually run — at 8k context, so it's a fair comparison. A card that can't fit Qwen3 235B-A22B (MoE)at that quality isn't listed, because it could only run a crushed version.

Best value

—

Best throughput per dollar among cards that fit.

Cheapest that runs it

—

Lowest-cost card that still fits the model.

No compromise

—

Most headroom and highest throughput.

Every card that runs it, ranked

Best quantisation each card fits at 8k context, with the cheapest live rental and a bandwidth-derived throughput estimate. Sorted by tokens/sec per dollar.

GPUVRAMVRAM used~tok/sCheapest rentaltok/s per $

How this ranking is made — and its limit

Fit and throughput are computed from Qwen3 235B-A22B (MoE)'s own configuration with the engine behind our cost-to-run page (bits-per-weight validated to 0.7% median error). Rental prices are the cheapest live figure across the providers we track, last updated Sun, 04 Oct 2026 18:21:15 GMT.

The value column ranks renting, because that's what we can price precisely. If you're buyinga card, the fit and throughput columns are exactly what you need — but we don't publish a purchase-price value ranking, because we don't have verified street prices and won't invent them. For the buy-vs-rent decision itself, see the guides below.

Tokens/sec is a bandwidth-derived estimate, not a benchmark — we don't run our own hardware tests (editorial policy).

Check it yourself

See exactly what fits, at any context length

Running Qwen3 235B-A22B (MoE): common questions

How much VRAM do you need to run Qwen3 235B-A22B (MoE)?
At Q4_K_M and 8k context, Qwen3 235B-A22B (MoE) needs about 135.0 GB — 132.7 GB of weights plus 1.5 GB of KV cache and ~0.8 GB overhead. So you want a card with at least that much free memory; the KV cache grows if you use longer context.
Can an RTX 4090 run Qwen3 235B-A22B (MoE)?
Not at Q4_K_M — Qwen3 235B-A22B (MoE) needs more than the card's 24 GB at that quality and 8k context. You'd have to drop to a lower quantisation (worse quality), shorten the context, or use a bigger card or two cards.
Can an RTX 3090 run Qwen3 235B-A22B (MoE)?
Not at Q4_K_M — Qwen3 235B-A22B (MoE) needs more than the card's 24 GB at that quality and 8k context. You'd have to drop to a lower quantisation (worse quality), shorten the context, or use a bigger card or two cards.
Can an RTX 4060 Ti 16GB run Qwen3 235B-A22B (MoE)?
Not at Q4_K_M — Qwen3 235B-A22B (MoE) needs more than the card's 16 GB at that quality and 8k context. You'd have to drop to a lower quantisation (worse quality), shorten the context, or use a bigger card or two cards.

Before you buy

Best GPU for other models