Gemma 4 26B-A4B locally: the Gemma that runs fast on 16GB

Gemma 4's mixture-of-experts model holds about 26B parameters but activates only 3.8B — so it's quick, and it fits a 16GB card at Q4. The VRAM math,…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

Gemma 4 26B-A4B is Google's mixture-of-experts Gemma 4: about 26.5B parameters total (our VRAM engine's count of the full model) but only 3.8B active per token, so it runs fast. At Q4_K_M our VRAM…

~26.5B total (card headlines 25.2B) / 3.8B active (MoE), text + image, 256K context, Apache-2.0 per the card.

Q4_K_M ~16.1 GB, Q8_0 ~27.4 GB (computed by our VRAM engine, 16K context).

Fits a 16GB card at Q4; run gemma4:26b, not the bare tag.

Aliteq

Read the full story

Gemma 4 26B-A4B locally: the Gemma that runs fast on 16GB

Read the full story on Aliteq