Gemma 4 31B locally: the 24GB-card Gemma

Gemma 4 31B is the big dense Gemma — the one that asks for a 24GB card and rewards it. The VRAM per quant, the GPU that fits, and how to run it…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

Gemma 4 31B is Google's large dense Gemma 4 (about 32.7B parameters as our VRAM engine counts the full model, multimodal, 256K context). At Q4_K_M our VRAM engine computes about 20.5 GB, so it fits…

~32.7B params (Google's card headlines 30.7B for the core model), dense, text + image, 256K context, Apache-2.0 per the card.

Q4_K_M ~20.5 GB, Q8_0 ~34.4 GB, BF16 ~63 GB (computed by our VRAM engine, 16K context).

Fits a 24GB card at Q4; run gemma4:31b, not the bare tag.

Aliteq

Read the full story

Gemma 4 31B locally: the 24GB-card Gemma

Read the full story on Aliteq