three free AI models all claim a 128K memory — only one actually holds onto it

Phi-4 Mini, Gemma 3, and Llama 3.2 all advertise roughly the same 128K context window. Benchmark data says that number means something very…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

Phi-4 Mini (3.8B), Gemma 3 (4B), and Llama 3.2 (3B) all advertise roughly 128K-131K token context windows.

The short version

Phi-4 Mini beats Llama 3.2 3B on raw accuracy — 73% vs 65% on MMLU, 62% vs 48% on MATH — despite being a similar size.

The short version

NVIDIA's RULER benchmark research found most models degrade meaningfully 30-40% before their claimed context limit — a 128K model is often shaky past roughly 80K-90K tokens.

The short version

All three run comfortably on 3-4GB of VRAM at Q4 quantization, so the choice rarely changes your hardware budget.

The short version

Gemma 3 is the only one of the three built natively multimodal — it reads images and screenshots without a separate vision pipeline.

Which one should you run?

For pure reasoning over long context on a budget GPU, Phi-4 Mini is the strongest of the three right now — the MMLU and MATH gap over Llama 3.2 is large enough to matter, and it's the smallest to…

Aliteq

Read the full story

three free AI models all claim a 128K memory — only one actually holds onto it

Read the full story on Aliteq