Bonsai 2 27B: How PrismML Shrinks a 27B Model to 5.9 GB
A 27B model that fits in about 6 GB sounds like a trick. It is ternary weights, a shared scale and a rotation. Here is how it works, what PrismML's…
Aliteq
Tensor · Local AI & Automation Editor
The short answer
Ternary Bonsai 2 27B is Qwen3.8 27B with almost every weight rounded to one of three values: -1, 0 or +1, plus a shared scale. That takes the language model from 53.8 GB to about 5.9 GB, roughly 9…
How it shrinks: three-value weights plus one 16-bit scale per 128 weights come to about 1.7 bits per weight, down from 16
What PrismML reports: 83.9 vs 85.4 average over 20 benchmarks, run by PrismML, not independently
Where it loses: SWE-bench Verified 60.8 vs 80.6 and Terminal-Bench 2.1 52.8 vs 69.7
The catch: stock llama.cpp cannot run the files; you need PrismML's fork
Aliteq
Read the full story
Bonsai 2 27B: How PrismML Shrinks a 27B Model to 5.9 GB