Adding CPU offload should make inference slower. Sometimes it makes it five times faster. Both are true, and the reason matters more than the trick.
The short answer
Why this particular split is clever
Aliteq
The MoE offload trick that makes your model 5× faster — and why it probably won't work for you