Best MoE Models for a 12GB or 16GB GPU: The --n-cpu-moe Number and the RAM Each Needs

Seven mixture-of-experts models, from gpt-oss 20B to Qwen3 235B, split into what stays on your card and what moves to system RAM. The card barely…

Aliteq
Voltage · Hardware Editor

The short answer

On a 12 GB or 16 GB card, the MoE model you can run is set by your system RAM, not your GPU. With 16 to 32 GB of RAM, run Qwen3.6 35B-A3B, Qwen3 30B-A3B or gpt-oss 20B. With 48 GB, add Qwen3-Next…

Starting values at 32K context: gpt-oss 120B is --n-cpu-moe 33 on 12 GB and 31 on 16 GB; Qwen3 30B-A3B is 31 and 20

The GPU part barely moves: every model keeps 11 to 14 GB on a 16 GB card; the rest is expert weights in RAM

Active size decides the feel: the A3B models read under 1 GB of experts from RAM per token, GLM-4.5-Air over 3 GB

Derived, not measured: we split the real GGUF files and ran our VRAM formula; we did not benchmark speed

Aliteq

Read the full story

Best MoE Models for a 12GB or 16GB GPU: The --n-cpu-moe Number and the RAM Each Needs

Read the full story on Aliteq