can you run gpt-oss-20b locally? Yes — on a single 16GB GPU (with one catch)

OpenAI's gpt-oss-20b fits on a 16GB graphics card and flies at 225 tokens/second on an RTX 4090. But there's a context-length trap that quietly eats…

Aliteq
Lena Fischer · AI & Local Compute Editor

What you need to run gpt-oss-20b

Fits on a single 16GB GPU — ~12.7GB VRAM at MXFP4 with 8K context.

What you need to run gpt-oss-20b

Fast: ~225 tokens/second on an RTX 4090 (8K context).

What you need to run gpt-oss-20b

The 128K-context trap: long context grows the KV cache and eats VRAM — 16GB fills up.

What you need to run gpt-oss-20b

Safe pick: 24GB if you want big context; 16GB is fine for everyday use.

What you need to run gpt-oss-20b

Also runs on a 24GB Mac (unified memory) — see best Mac for local AI.

What you need to run gpt-oss-20b

Easiest way: pull it in Ollama or LM Studio.

Aliteq

Read the full story

can you run gpt-oss-20b locally? Yes — on a single 16GB GPU (with one catch)

Read the full story on Aliteq