Best GPU for Running gpt-oss Locally in 2026 (20B and 120B)

OpenAI shipped two very different models. The 20B fits a normal 16GB card and flies; the 120B is an 80GB-class model you'll mostly rent. Here's the…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

gpt-oss-20B is the easy one. It loads in ~14GB at its native MXFP4 precision, so 16GB is my practical minimum and 24GB is comfortable. Reviewers report ~225 tok/s on an RTX 4090 — it flies.

The short version

Best value for the 20B: any 16GB card (RX 9070 / RTX 5070 Ti / RTX 5060 Ti 16GB) or, my pick on value, a used RTX 3090 (24GB) for headroom.

The short version

gpt-oss-120B is an 80GB-class model. The ~60.8GiB MXFP4 checkpoint runs on a single 80GB GPU per OpenAI; ~60GB is a constrained floor, not comfort.

The short version

For the 120B, I'd rent, not buy. An 80GB H100 by the hour beats a five-figure card that idles — unless you run it daily.

The short version

Or go unified memory (a 128GB Ryzen AI / Strix Halo box or a big Mac) to hold the 120B slowly on a budget.

What I'd run gpt-oss on

For gpt-oss-20B: a 16GB card is the floor and a used RTX 3090 (24GB) is my value pick — it's fast on anything modern, so buy VRAM, not speed. For gpt-oss-120B: treat it as an 80GB-class model — rent…

Aliteq

Read the full story

Best GPU for Running gpt-oss Locally in 2026 (20B and 120B)

Read the full story on Aliteq