The GitHub request to keep MoE experts in RAM is still open after 14 months. But since Ollama 0.30, llama.cpp's own fit step can park them there for you, as long as you leave num_gpu alone.
You read about --n-cpu-moe, the llama.cpp flag that parks a mixture-of-experts model's bulky expert weights in system RAM. Then you went looking for the same switch in Ollama. It isn't there. Not in the Modelfile docs, not in the API options, not in the environment variables.
That part of the story is true. The rest of it changed in May 2026, and most of the advice floating around hasn't caught up. I read Ollama's source and llama.cpp's source on 4 October 2026 to work out what actually happens to your experts today. I'm a spreadsheet person, not a Go developer, so I'll show you exactly which lines I'm leaning on.
Can Ollama keep MoE experts on the CPU?
Not on request. Ollama has no option that says "put the experts of the first N layers in RAM". The feature request is still open, and Ollama's documented knobs don't include it. What Ollama does have, since version 0.30, is llama.cpp underneath, and llama.cpp's automatic placement already knows what an expert is.
Here are the options Ollama accepts when it loads a model, read in its api/types.go: num_ctx, num_batch, num_gpu, main_gpu, use_mmap, num_thread and draft_num_predict. That's the whole list. None of them mentions experts. The same list feeds Modelfile PARAMETER lines, so there's no hidden Modelfile key either.
Compare that with the two other tools most people run. llama.cpp has had --n-cpu-moe since 4 August 2025 (PR #15077). LM Studio added a "Force Model Expert Weights onto CPU" toggle in version 0.3.23, eight days later. If you want the mechanics of the flag itself, our explainer on what --n-cpu-moe actually does covers it.
Where the GitHub request stands
Issue #11772 is open, labeled "feature request", with 41 comments. It was opened on 7 August 2025, and its last activity was 5 June 2026. No setting has shipped. The one clear statement of intent from the Ollama side came on a related pull request in September 2025.
The issue's title is "use cpu to offload moe weights to reduce the VRAM usage." Its first post proposed a strategy setting with three values: FULL, PARTIAL and NONE. Most of the 41 comments are votes, user benchmarks and patches.
A contributor did write a pull request, #12333, "feat: add support for MoE offloading", opened 18 September 2025. It is still open. The reply it got the same day, from Ollama contributor jessegross, explains why:
"Thank you but we want to be able to configure this automatically based on available memory rather than making the user configure it, similar to how the rest of Ollama works."
jessegross on ollama/ollama PR #12333, 18 Sep 2025
The same reply said they were "generally" not adding new features to the old llama engine. Keep that sentence in mind. Ollama later replaced that engine, and the replacement does configure this automatically.
Dates from GitHub and the LM Studio release notes, read 4 Oct 2026. · aliteq research
What Ollama does with an MoE model today
Since 0.30, Ollama starts llama.cpp's llama-server for every GGUF model and passes it a short list of flags. Whether it passes a GPU layer count depends only on num_gpu. With num_gpu unset, it passes none, and llama.cpp's fit step, which is on by default, decides where every tensor goes.
The 0.30 release note (13 May 2026) put it simply: "Ollama 0.30 is now available, with improved compatibility and performance using llama.cpp." The source is blunter. A comment in llm/server.go reads: "All GGUF models are served via the upstream llama-server subprocess."
Then llm/llama_server.go translates num_gpu into a llama.cpp flag. This is the table I care about:
What Ollama's num_gpu becomes in llama-server (Ollama v0.35.1, read 4 Oct 2026)
num_gpu not set (default)
Flag Ollama passes
None ("let llama-server auto-detect")
Who places the weights
llama.cpp's fit step, on by default
num_gpu 0
Flag Ollama passes
-ngl 0
Who places the weights
You: CPU only
num_gpu N (any positive number)
Flag Ollama passes
-ngl N
Who places the weights
You: N whole layers in VRAM
Flag Ollama passes
Who places the weights
num_gpu not set (default)
None ("let llama-server auto-detect")
llama.cpp's fit step, on by default
num_gpu 0
-ngl 0
You: CPU only
num_gpu N (any positive number)
-ngl N
You: N whole layers in VRAM
Now the llama.cpp side, read in common/fit.cpp at build b11232, the build Ollama 0.35.1 pins. The fit step only touches settings you didn't set. For an MoE model that doesn't fit, it first measures memory "with all MoE tensors moved to system memory". It fills the GPU with the small, always-used parts. Then it moves expert weights back onto the GPU, front to back, until the card is full.
That is --n-cpu-moe, chosen for you. It's also roughly what the September 2025 reply asked for: placement "based on available memory rather than making the user configure it". I can't tell you whether Ollama's team sees it that way. The issue is still open, and nobody on it says so.
One Ollama detail matters here. Ollama always passes the context size (-c) itself. The fit step leaves a user-set context alone. So it can't shrink your context to make room. It can only move weights. A big context means more experts in RAM.
Our reading of Ollama's and llama.cpp's source, 4 Oct 2026. Not a benchmark. · aliteq research
The trap: setting num_gpu turns the smart part off
If you set num_gpu to any number, llama.cpp's fit step treats the layer count as yours and stops placing weights. You then get a plain split by whole layers: attention and experts move together. That is the old way of running a too-big model, and it's exactly what the expert-aware split was built to beat.
The line in fit.cpp is unambiguous. If the GPU layer count differs from the default, it stops with "n_gpu_layers already set by user ... abort". The log then shows "failed to fit params to free device memory".
That's why the folk advice "raise num_gpu until it fits" deserves a second look on Ollama 0.30 and later. On the older engine it was the only lever you had. On the current one, it switches off the placement that knows about experts. The step-by-step guide to running a big MoE model on a small GPU explains why attention belongs on the GPU and experts in RAM, and not the other way round.
So, if you set num_gpu in a Modelfile or an API call for an MoE model, try removing it first. Then compare.
How to check what Ollama actually did
Two places tell you. ollama ps shows the overall split between CPU and GPU memory. The Ollama server log carries llama-server's own startup lines, because Ollama runs it with verbose logging on purpose. Look there for the layer count and for any "failed to fit" warning.
**Run `ollama ps`.** The PROCESSOR column reads like "48%/52% CPU/GPU" when a model sits partly in system memory, per Ollama's FAQ. For an MoE model with experts in RAM, a split here is expected, not a failure.
**Open the server log.** Ollama starts llama-server with `--log-verbosity 4`, with a source comment saying it keeps "startup memory/offload lines visible". Look for "offloaded N/M layers to GPU".
**Look for "failed to fit params to free device memory".** It means something you set (num_gpu, or an override) took placement away from the fit step.
**Shrink num_ctx if experts crowd out.** Ollama's default context is 4K, 32K or 256K depending on your VRAM. A smaller window leaves more VRAM for experts.
A word of caution about ollama ps. Issue #15237 (April 2026) reported Gemma 4 showing "100% GPU" while actually running on the CPU. That bug was closed as completed on 7 April 2026. Release 0.30.11 also fixed ollama ps double-counting memory-mapped weights on partial offload. On a current version the column should be honest, but the log is the ground truth.
Workaround 1: the inherited environment variable (unofficial)
Ollama launches llama-server with a copy of its own environment. llama.cpp reads LLAMA_ARG_N_CPU_MOE as the environment form of --n-cpu-moe. Put the two together and a variable set on the Ollama server reaches the engine. It works only because of how the code is written today, and Ollama documents none of it.
The source facts: Ollama's SetupLlamaServerCommandEnv starts from os.Environ(), the server's full environment. In llama.cpp b11232, --n-cpu-moe is tied to LLAMA_ARG_N_CPU_MOE and --cpu-moe to LLAMA_ARG_CPU_MOE. Ollama does document two llama.cpp variables of the same family, LLAMA_ARG_FIT and LLAMA_ARG_FIT_TARGET, in ollama serve help. The MoE ones aren't on that list. One user on issue #11772 reported using LLAMA_ARG_CPU_MOE this way in June 2026. We have not tried it.
If you try it, set the variable where Ollama's FAQ says server variables go: launchctl setenv for the Mac app, systemctl edit ollama.service with an Environment= line on Linux, or your user environment variables on Windows. Then restart Ollama.
Workaround 2: use llama.cpp or LM Studio for this model
If you want a per-model expert count that you control, use the tools that expose one. llama.cpp gives you --n-cpu-moe directly. LM Studio gives you a per-layer slider in a desktop app (an all-or-nothing toggle before 0.4.0). Keep Ollama for the models that fit, and run the big MoE one where you can set the number.
llama.cpp.llama-server -m model.gguf --n-cpu-moe N is the exact control the Ollama issue asks for. It also speaks the OpenAI API format, so most apps that talk to Ollama can point at it instead. The llama.cpp MoE offloading guide covers the flags and the -ot patterns. The gpt-oss-20b on an 8GB or 12GB card walkthrough applies it to one popular model.
LM Studio. Since 0.3.23, its advanced load settings include "Force Model Expert Weights onto CPU". Its release notes describe it as "the same underlying technology as llama.cpp's --n-cpu-moe" and add a fair warning: "If you can offload the entire model to GPU memory, you're better off sticking with placing expert weights onto the GPU as well". Settings can be saved per model. If you're choosing between the two apps, our Ollama vs LM Studio comparison covers the rest of the trade.
Quick answers
Does Ollama support --n-cpu-moe?
Not as an Ollama setting. Its load options are num_ctx, num_batch, num_gpu, main_gpu, use_mmap, num_thread and draft_num_predict, and none controls experts. Issue #11772 asking for it has been open since 7 August 2025. Since Ollama 0.30, though, llama.cpp's automatic fit step can put experts in RAM when num_gpu is left unset.
Is there an OLLAMA_MOE_OFFLOAD environment variable?
No. OLLAMA_MOE_OFFLOAD with FULL, PARTIAL or NONE was a proposal in the first post of issue #11772. It was never added. Ollama's documented server variables include LLAMA_ARG_FIT and LLAMA_ARG_FIT_TARGET, but no MoE variable.
Should I set num_gpu for a big MoE model in Ollama?
On Ollama 0.30 or later, try leaving it unset first. Any num_gpu value makes llama.cpp's fit step stand aside, and you get a split by whole layers instead of the expert-aware one. num_gpu 0 forces the whole model onto the CPU.
Does the LLAMA_ARG_N_CPU_MOE trick work in Ollama?
The source suggests it can. Ollama hands its environment to llama-server, and llama.cpp reads that variable as --n-cpu-moe. Ollama doesn't document it, it applies to every model on that server, and we haven't tested it. Treat it as unsupported.
Why does ollama ps show a CPU/GPU split for my MoE model?
Because part of the model sits in system memory. For an MoE model on a small card that's expected: the experts that don't fit live in RAM. Check the server log for the layer count and for any "failed to fit" warning to see who made the split.
Does any of this apply to older Ollama versions?
No. Before 0.30, Ollama ran its own engine, and the llama.cpp fit behavior described here wasn't in the path. Update to a current release, or use llama.cpp or LM Studio for MoE offload.