
AI
how to run a 30B model on a 12GB GPU: the one llama.cpp flag that makes it possible
--n-cpu-moe is the setting that lets a modest graphics card punch far above its VRAM. It offloads the parts of a Mixture-of-Experts model you use least to system RAM — and here's exactly how to tune it for your card.
Lena Fischer · 2h ago · 9 min