The best GPU for running Qwen3.5 122B-A10B (MoE) (2026)
Ranked from live rental prices and computed VRAM fit, not opinion. Every card is judged at Q4_K_M — the quantisation most people actually run — at 8k context, so it's a fair comparison. A card that can't fit Qwen3.5 122B-A10B (MoE)at that quality isn't listed, because it could only run a crushed version.
Best value
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Best throughput per dollar among cards that fit.
Cheapest that runs it
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Lowest hourly rental that fits it — $1.690/hr, Q4_K_M.
No compromise
—
Most headroom and highest throughput.
Every card that runs it, ranked
Best quantisation each card fits at 8k context, with the cheapest live rental and a bandwidth-derived throughput estimate. Sorted by tokens/sec per dollar.
| GPU | VRAM | VRAM used | ~tok/s | Cheapest rental | tok/s per $ |
|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | 72 GB75% | — | $1.690 | — |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | 72 GB75% | — | — | — |
How this ranking is made — and its limit
Fit and throughput are computed from Qwen3.5 122B-A10B (MoE)'s own configuration with the engine behind our cost-to-run page (bits-per-weight validated to 0.7% median error). Rental prices are the cheapest live figure across the providers we track, last updated Sun, 04 Oct 2026 19:21:04 GMT.
The value column ranks renting, because that's what we can price precisely. If you're buyinga card, the fit and throughput columns are exactly what you need — but we don't publish a purchase-price value ranking, because we don't have verified street prices and won't invent them. For the buy-vs-rent decision itself, see the guides below.
Tokens/sec is a bandwidth-derived estimate, not a benchmark — we don't run our own hardware tests (editorial policy).
Check it yourself
See exactly what fits, at any context length
Running Qwen3.5 122B-A10B (MoE): common questions
How much VRAM do you need to run Qwen3.5 122B-A10B (MoE)?
Can an RTX 4090 run Qwen3.5 122B-A10B (MoE)?
Can an RTX 3090 run Qwen3.5 122B-A10B (MoE)?
Can an RTX 4060 Ti 16GB run Qwen3.5 122B-A10B (MoE)?
What's the cheapest way to run Qwen3.5 122B-A10B (MoE)?
Before you buy
GLM-4.5-Air on local hardware
The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.
Cheapest way to run GLM in the cloud
Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.
Run Mistral Small 24B locally
VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.
Cheapest RTX 4090 rental
Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.