which quantization should you use? Q4 vs Q5 vs Q8, settled with real numbers

Every model comes in a dozen quant sizes and it's paralyzing. Here's the simple truth: Q4_K_M is the default for almost everyone, and here's exactly…

Aliteq
Lena Fischer · AI & Local Compute Editor

The pick, in one glance

Q4_K_M — the default. ~3.5% quality loss, 75% smaller than FP16. Right for almost everyone.

The pick, in one glance

Q5_K_M — worth it with 12GB+ VRAM, especially for coding and complex reasoning.

The pick, in one glance

Q8_0 — overkill. Near-FP16 quality but heavy and slow; the gain is imperceptible in normal chat.

The pick, in one glance

Always prefer K_M over K_S — it protects the attention layers that matter most, for a tiny size bump.

The pick, in one glance

Avoid Q2/Q3 unless desperate to fit — that's where quality really degrades.

The pick, in one glance

Rule: pick the highest quant that leaves VRAM headroom; Q4_K_M if unsure.

Aliteq

Read the full story

which quantization should you use? Q4 vs Q5 vs Q8, settled with real numbers

Read the full story on Aliteq