A 27B model that fits in about 6 GB sounds like a trick. It is ternary weights, a shared scale and a rotation. Here is how it works, what PrismML's own tests show, and where the loss hides.
I spend a lot of time shrinking models, and most claims of "a 27B model on a laptop" fall apart on the second read. This one mostly holds up, and the reasons are interesting. The details below come from PrismML's launch post, its whitepaper and its Hugging Face model card. I read them on 3 October 2026. I have not run the model, and every score is PrismML's own.
What Bonsai 2 27B actually is
Ternary Bonsai 2 27B is PrismML's compressed version of Qwen3.8 27B. It is not a new model trained from scratch. The whitepaper says it uses the same hybrid-attention architecture as the base, and the weights are Apache 2.0 licensed.
PrismML's launch post gives a 5.9 GB total footprint and a 262K-token context window. It also takes images. The earlier Bonsai 27B came out in July 2026, and this release swaps in the newer Qwen3.8 base.
Two details trip people up. The model is ternary, not 1-bit. PrismML also sells a 1-bit sibling from the first generation, so read the name carefully. And 5.9 GB covers the text model only. The vision part is a separate optional file of 0.63 GB.
How ternary compression works
Each weight becomes -1, 0 or +1, and a group of 128 weights shares one 16-bit scale. Three values need log2(3), about 1.585 bits. The shared scale adds 16 / 128 = 0.125 bits per weight. That totals about 1.71 bits, against 16 in the original.
Bit arithmetic is ours; sizes from PrismML's whitepaper, Table 3. · aliteq research
Rounding every weight to three values would normally wreck a model. PrismML adds two things to limit the damage. First, it stores the weights in a rotated basis, using a Hadamard rotation in blocks of 1,024. The runtime applies the matching rotation to the activations, so the math comes out the same. PrismML says the rotation costs no extra bits.
Second, it keeps a few delicate tensors at full precision. These are the recurrent-state tensors of the linear-attention layers and the normalization weights. They total 26.2 million parameters, 0.0976% of the language model. Everything else is ternary, including the embeddings and the output layer.
Why the file is 5.93 GB, not 5.8
The ideal ternary size is 5.8 GB. Real files need a layout a GPU can read fast. PrismML ships two. PTQ1_0 packs the three-value codes densely and lands at 1.76 bits per weight, 5.93 GB in the whitepaper. PQ2_0 gives each code a 2-bit slot, which wastes space but unpacks faster: 2.16 bits per weight, 7.25 GB.
Neither is simply better. PrismML says PTQ1_0 decodes faster on Ada-generation cards such as the RTX 4090 and on the L4. PQ2_0 is faster on H100, A100 and Blackwell cards, and for reading long prompts everywhere. The Hugging Face card lists 5.95 and 7.21 GB for the same two files, so expect small differences between PrismML's pages.
PrismML whitepaper, Tables 3 and 7 and section 3. The 9.1x ratio is our division: 53.8 / 5.93. · aliteq research
What PrismML says it keeps
PrismML reports an average of 83.9 across 20 benchmarks, against 85.4 for full-precision Qwen3.8 27B. That is 98.2%, a gap of 1.5 points. It also reports a conventional 2-bit build (IQ2_XXS, 7.3 GB) at 75.2 on the same suite.
The Hugging Face card uses a different, smaller suite of 14 benchmarks: 84.78 vs 86.32. The percentage is the same, 98.2%, but the scores are not comparable with the launch post. I quote the whitepaper suite throughout.
Bonsai 2 27B vs full-precision Qwen3.8 27B (PrismML's runs)
Overall (20 benchmarks)
Bonsai 2 27B
83.9
Qwen3.8 27B
85.4
Math
Bonsai 2 27B
96.57
Qwen3.8 27B
97.06
Coding (short problems)
Bonsai 2 27B
81.58
Qwen3.8 27B
82.17
Instruction following
Bonsai 2 27B
82.66
Qwen3.8 27B
81.25
Tool calling
Bonsai 2 27B
77.57
Qwen3.8 27B
79.74
Knowledge and reasoning
Bonsai 2 27B
83.95
Qwen3.8 27B
86.66
Vision
Bonsai 2 27B
78.59
Qwen3.8 27B
81.64
Bonsai 2 27B
Qwen3.8 27B
Overall (20 benchmarks)
83.9
85.4
Math
96.57
97.06
Coding (short problems)
81.58
82.17
Instruction following
82.66
81.25
Tool calling
77.57
79.74
Knowledge and reasoning
83.95
86.66
Vision
78.59
81.64
Math, short coding and instruction following are nearly level. The losses concentrate in knowledge and vision. That fits the intuition that rounding weights hurts recall of facts more than it hurts step-by-step reasoning.
Where the 98.2% hides a bigger gap
Long coding-agent tasks lose about a quarter of the score. In the whitepaper, SWE-bench Verified is 60.8 for Bonsai against 80.6 for the full model. Terminal-Bench 2.1 is 52.8 against 69.7. That is 75.4% and 75.8% retained, by our division.
Bonsai score divided by Qwen3.8 27B FP16 score, from PrismML's whitepaper. Our arithmetic. · aliteq research
PrismML does not hide this. The whitepaper itself says the model keeps "roughly three quarters" on both. Still, the headline number is an average over many short tests. If you want an agent that works through a repository for an hour, the 75% figure is the one to plan around. Errors compound over many steps.
The known-issues page adds a practical warning. PrismML lists malformed or looping tool calls as a known limitation in structured output, still open when it last checked on 23 September.
What it takes to run it
You need PrismML's own llama.cpp build. The model card is blunt: stock llama.cpp rejects the two file types as unknown, and it loads a similar Q2_0 type without warning and produces garbage, because it lacks the rotation step. That also rules out tools built on stock llama.cpp until they adopt the fork. PrismML's launch post names CUDA on NVIDIA GPUs and MLX on Apple devices.
It is also a reasoning model that thinks by default. The card warns that a small output cap gives empty or cut-off answers, and recommends a limit of 16,384 tokens or more. If you serve models to your own apps, our guide to running local AI as an API server covers the basics, though this model needs the fork's binary.
Will it fit your card?
Weights are only part of the memory bill. By our formula, the 5.93 GB file is 5.52 GiB. Add the conversation cache and 0.78 GiB of runtime overhead. Qwen3.8 27B has 16 full-attention layers, 4 key-value heads and a head size of 256, which I read from its config. That makes the cache 2 x 16 x 4 x 256 x 2 bytes = 64 KiB per token.
At 32K tokens that totals 5.52 + 2.00 + 0.78 = 8.3 GiB. At 128K tokens it is 14.3 GiB. This assumes the cache is stored at 16-bit and Bonsai keeps the base architecture, which the whitepaper states. The optional vision file adds 0.63 GB. PrismML lists sub-4-bit cache compression only as early results, so do not count on it.
PrismML's whitepaper reports 142.5 tokens per second on an RTX 5090 and 90.9 on an RTX 4090 for PQ2_0. On Apple chips it reports 46.8 on an M5 Max, 27.7 on an M5 Pro and 18.0 on an M4 Pro. The M4 figure comes from an earlier build. The Hugging Face card shows lower numbers for the same cards, for example 129.9 on the 5090, so treat speeds as a range. For model-by-model memory math, use our cost to run tool and the best GPU for local AI guide.
Should you use it?
It depends on the job. For chat, summaries, math help and short code edits, PrismML's numbers say you give up little against the full model. For a long autonomous coding agent, the data says you give up a quarter. If you have 24 GB or more, a Q4 build of the full model is a safer baseline, and PrismML's own card puts a 4-bit build at 85.18 against Bonsai's 84.78 on its 14-test suite, at about three times the size.
The honest verdict is that the compression is real, the memory saving is large, and independent tests are still missing. Wait for someone other than PrismML to reproduce the scores before you build on them.
Bonsai 2 27B: quick answers
Is Bonsai 2 27B a 1-bit model?
No. It is ternary: weights are -1, 0 or +1, with a shared 16-bit scale per 128 weights, about 1.7 bits per weight. PrismML's first-generation Bonsai 27B also had a 1-bit version, which is a different model.
Is it really 5.9 GB?
The text model is 5.93 GB in the whitepaper and 5.95 GB on the Hugging Face card. That excludes the optional 0.63 GB vision file, and the cache and runtime need extra memory on top.
Does it keep 98.2% of Qwen3.8 27B's quality?
PrismML reports 83.9 vs 85.4 across 20 benchmarks, which is 98.2%. PrismML ran those tests itself, and on SWE-bench Verified and Terminal-Bench 2.1 it keeps only about 75%.
Can I run it in stock llama.cpp?
No. The model card says stock llama.cpp rejects the file types or produces garbage. Use PrismML's llama.cpp fork, or its MLX build on Apple hardware.
What is the licence?
Apache 2.0, per PrismML's launch post and the Hugging Face card. The base model, Qwen3.8 27B, is also listed as Apache 2.0.
What is the difference between PTQ1_0 and PQ2_0?
PTQ1_0 packs the three-value codes densely, 5.93 GB. PQ2_0 gives each code a 2-bit slot, 7.25 GB. PrismML says PTQ1_0 decodes faster on Ada-generation cards and the L4, and PQ2_0 on H100, A100 and Blackwell cards.