Ollama has no --n-cpu-moe switch. Here's what it does with MoE experts instead

The GitHub request to keep MoE experts in RAM is still open after 14 months. But since Ollama 0.30, llama.cpp's own fit step can park them there for…

Aliteq
Voltage · Hardware Editor

The short answer

No, Ollama has no setting like llama.cpp's --n-cpu-moe. The request, issue 11772, has been open since 7 August 2025. But since Ollama 0.30 (13 May 2026), every GGUF model runs inside llama.cpp's own…

Leave num_gpu alone. Any value you set counts as a manual choice, and the expert-aware fit step steps aside.

There is an unofficial route. Ollama passes its own environment to llama-server, so LLAMA_ARG_N_CPU_MOE reaches the engine. Ollama doesn't document it, and it applies to every model.

Want a per-model number? Run llama.cpp directly, or use LM Studio's expert-weights toggle.

We read code, we didn't benchmark. No speed numbers here are ours.

Read this before you set it

The variable is server-wide. It applies the same N to every model that server loads, including dense models (where it does nothing) and small MoE models that would have fit entirely on the GPU…

Aliteq

Read the full story

Ollama has no --n-cpu-moe switch. Here's what it does with MoE experts instead

Read the full story on Aliteq