aliteq.

The best GPU for running gpt-oss 20B (MoE) (2026)

Ranked from live rental prices and computed VRAM fit, not opinion. Every card is judged at Q4_K_M — the quantisation most people actually run — at 8k context, so it's a fair comparison. A card that can't fit gpt-oss 20B (MoE)at that quality isn't listed, because it could only run a crushed version.

Best value

NVIDIA RTX A4000

Best throughput per dollar among cards that fit.

Cheapest that runs it

NVIDIA RTX A4000

Lowest hourly rental that fits it — $0.066/hr, Q4_K_M.

No compromise

—

Most headroom and highest throughput.

Every card that runs it, ranked

Best quantisation each card fits at 8k context, with the cheapest live rental and a bandwidth-derived throughput estimate. Sorted by tokens/sec per dollar.

GPUVRAMVRAM used~tok/sCheapest rentaltok/s per $
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition96 GB13 GB14%—$1.690—
NVIDIA RTX PRO 6000 Blackwell Workstation Edition96 GB13 GB14%———
NVIDIA RTX 6000 Ada Generation48 GB13 GB27%—$0.494—
AMD Radeon PRO W790048 GB13 GB27%———
NVIDIA RTX A600048 GB13 GB27%—$0.330—
NVIDIA GeForce RTX 509032 GB13 GB41%—$0.402—
NVIDIA RTX 5000 Ada Generation32 GB13 GB41%—$0.336—
NVIDIA GeForce RTX 409024 GB13 GB54%—$0.334—
AMD Radeon RX 7900 XTX24 GB13 GB54%———
NVIDIA GeForce RTX 3090 Ti24 GB13 GB54%—$0.120—
NVIDIA RTX A500024 GB13 GB54%—$0.160—
NVIDIA GeForce RTX 309024 GB13 GB54%———

How this ranking is made — and its limit

Fit and throughput are computed from gpt-oss 20B (MoE)'s own configuration with the engine behind our cost-to-run page (bits-per-weight validated to 0.7% median error). Rental prices are the cheapest live figure across the providers we track, last updated Sat, 10 Oct 2026 21:20:59 GMT.

The value column ranks renting, because that's what we can price precisely. If you're buyinga card, the fit and throughput columns are exactly what you need — but we don't publish a purchase-price value ranking, because we don't have verified street prices and won't invent them. For the buy-vs-rent decision itself, see the guides below.

Tokens/sec is a bandwidth-derived estimate, not a benchmark — we don't run our own hardware tests (editorial policy). This is a mixture-of-experts model without a known active-parameter count, so throughput is omitted rather than guessed.

Check it yourself

See exactly what fits, at any context length

Running gpt-oss 20B (MoE): common questions

How much VRAM do you need to run gpt-oss 20B (MoE)?
At Q4_K_M and 8k context, gpt-oss 20B (MoE) needs about 13.0 GB — 11.8 GB of weights plus 0.4 GB of KV cache and ~0.8 GB overhead. So you want a card with at least that much free memory; the KV cache grows if you use longer context.
Can an RTX 4090 run gpt-oss 20B (MoE)?
Yes. At Q4_K_M it uses about 13.0 GB of the card's 24 GB (54%).
Can an RTX 3090 run gpt-oss 20B (MoE)?
Yes. At Q4_K_M it uses about 13.0 GB of the card's 24 GB (54%).
Can an RTX 4060 Ti 16GB run gpt-oss 20B (MoE)?
Yes. At Q4_K_M it uses about 13.0 GB of the card's 16 GB (81%).
What's the cheapest way to run gpt-oss 20B (MoE)?
Among cards that fit it at Q4_K_M, the cheapest to rent right now is the NVIDIA RTX A4000 at about $0.066/hour. Whether renting or buying is cheaper overall depends on how many hours a day you'll actually use it.

Before you buy

Best GPU for other models