Bonsai 2 27B: How PrismML Shrinks a 27B Model to 5.9 GB

A 27B model that fits in about 6 GB sounds like a trick. It is ternary weights, a shared scale and a rotation. Here is how it works, what PrismML's…

Aliteq
Tensor · Local AI & Automation Editor

The short answer

Ternary Bonsai 2 27B is Qwen3.8 27B with almost every weight rounded to one of three values: -1, 0 or +1, plus a shared scale. That takes the language model from 53.8 GB to about 5.9 GB, roughly 9…

How it shrinks: three-value weights plus one 16-bit scale per 128 weights come to about 1.7 bits per weight, down from 16

What PrismML reports: 83.9 vs 85.4 average over 20 benchmarks, run by PrismML, not independently

Where it loses: SWE-bench Verified 60.8 vs 80.6 and Terminal-Bench 2.1 52.8 vs 69.7

The catch: stock llama.cpp cannot run the files; you need PrismML's fork

Aliteq

Read the full story

Bonsai 2 27B: How PrismML Shrinks a 27B Model to 5.9 GB

Read the full story on Aliteq