Gemma 4 26B-A4B locally: the Gemma that runs fast on 16GB
Gemma 4's mixture-of-experts model holds about 26B parameters but activates only 3.8B — so it's quick, and it fits a 16GB card at Q4. The VRAM math,…
Aliteq
Lena Fischer · AI & Local Compute Editor
The short answer
Gemma 4 26B-A4B is Google's mixture-of-experts Gemma 4: about 26.5B parameters total (our VRAM engine's count of the full model) but only 3.8B active per token, so it runs fast. At Q4_K_M our VRAM…
~26.5B total (card headlines 25.2B) / 3.8B active (MoE), text + image, 256K context, Apache-2.0 per the card.