Bonsai 27B: A 27B-Class model that runs on a phone
prismml.com
prismml.com
Good evaluation from 2024 https://arxiv.org/pdf/2402.18158
I'm currently working towards an updated version (not an og author), curious if others are aware of similar surveys, as I have yet to do a real lit search.
The 12B QAT model is indeed sort of mindblowing.
Hopefully more of the lab releases are trained under QAT so we can all benefit.
The best local agentic coding experience I’ve had so far is Qwen3.6-27B with Pi.
I have two main tasks I want to see if I can improve, coding and doc understanding/summary
Yes and no. I think where frontier models really blow small models away is in how thorough they are in order to infer your intentions and how best to accomplish them. So you can tell Claude "change this code to make it do X", whereas a Qwen3.6-27B or Gemma4-31B can do the same, but you have to be a lot more thorough, i.e., "change this code to make it do X, but first, let me explain the concept of X as I see it and some notes about things to avoid or pay attention to while you're doing it." So for best success with small models you really need a big toolbelt of skills and MCPs.
Imagine I had [a model that was good at math], [model that was good at code], [model that was good at writing], [model that was good at general knowledge]. If I then had [a model that was good at determining whether the user query would be best served by one of those models and sent it to it, leaving the rest of the models inactive], that is the platonic version of what MoE is. In practice, it works a bit differently. It instead basically restricts the number of pathways that can be utilized in solving problems during training, which allows for "expert neuron groupings" to form and "classifier layers" to form earlier on in the structure, but the effect is the same (better, even, since it allows some overlap between structures of experts). It also allows "routing to an expert" to happen token-by-token rather than at the prompt level.
2) Bitter lesson is misunderstood. Specialization and inductive biases still matter. ChatGPT isn’t the best chess player in the world just because it’s seen more math problems or read more Japanese poetry. Stockfish is, because it bakes in useful inductive biases like minimax.
In some workflows you might have 20 different models doing their specialized tasks. Pose detection, hand/eye/face detailers, classifiers, refiners, up scalers, taggers, etc can all use their own models and that’s not even the including the model(s) used for the actual image generation part.
I’m interested to see the optimization when this concept gets applied to other general ai tasks.
PHP/Wordpress code seems OK (better than the Gemma) but it gets stuck in reasoning loops.
Mind you, I am something of a cynic about the underlying 27B dense Qwen; I think the 35B MoE model is often better and it is just so, so much faster.
If these buddies are similarly bad on text, then they definitely don’t get anywhere close to big boys, no matter what the synthetic stats claim upon release.
If you can "hide" different models of 8GB VRAM requirements each that have those specialties and mix and match them for me without having to manage it manually, I'll be impressed. Until then I will keep using my Claude, because "remarkably good _for their size_" models I've tried so far just sucked at trying to use them the way I code at work with Claude.
Worse than Gemma at tool calling? Gemma's already bottom tier at that (at least when there's Qwen to compare to), that would just be unable to do tool calling at all.
* I will say that early on there were a LOT of issues with the chat template, across all engines. I dunno who decided using crappy Jinja templates was a good idea, but clearly it has its limitations. In the latest version of vLLM (0.25) they've ditched the Jinja templates for an in-engine parser and I've seen no issues.
But, what I really want is for Google to release bigger Gemma 4 models, particularly a bigger MoE, like a ~70B or ~120B. Gemma 4 is the best all-rounder among the models I can self-host even though I've got a 128GB Strix Halo. A 4-bit QAT version of a 70B MoE would probably be the sweet spot.
A bigger Qwen 3.6 with a 4-bit QAT version would also be welcome, as the prior bigger versions aren't notably better than 3.6 27B, but I guess Qwen is done doing larger open weights models. They did release AgentWorld recently, a post-train of the 3.6 MoE, so they're still doing some open things.
I think I want to see more third-party testing of this ternary Qwen to know if crushing it to 1.56 bits kills it; there are tons of benchmarks of Qwen 3.6 27B, so it's an ideal candidate to figure out what the extreme compression does to it.
How does this model compare to a recent 4G model? How do we know it retained intelligence from the parent rather then being fine tuned for the benchmarks?
I am not shtng on them or anything. I'd rather find it amazing, BUT given my limited knowledge, I feel the results miss fair comparison plots and the ones might be misleading. Buy I also reckon it might be me the problem. Anyone care to explain this poor silly fellow some of those points?
They don't give a F about AI or any new AI model that was announced this morning. Wasn't there news a while ago about them buying Perplexity?
I think their strategy is broadly correct, actually; I think there's still a bit of scope for "sit and wait and do it right" here. But acquiring more edge AI tech and edge AI people would potentially be in their interest.
Whether things are different now I don't know, but Prism would potentially be a good acquisition.
This is not the first Bonsai model from Prism using this technology, and they've also applied this technology to an image model.
The model aspect isn't that significant (even in this news), because it is Qwen, under the hood. The encoding efficiency is.
Prism's technology might be valuable, and their team could well be.
And they very evidently do care a lot about efficient on-device AI; they just don't care about developing frontier cloud models.
They could more or less redistribute Gemma as-is in the developer program; they are unlikely to be troubled by any of the licence terms.
https://www.theregister.com/on-prem/2000/08/02/jobs-snubs-at...
apple’s secrecy agenda has been defeated to an extent by the practicalities of ubiquitous technology?
I've tried a couple in LM Studio - the GGUF one and the MLX one - but neither worked there. Anyone else get them to work? Might be that LM Studio needs to upgrade their llama.cpp or MLX engines first.
Details are here -> https://github.com/PrismML-Eng/Bonsai-demo/blob/main/README....
You can also join Discord to communicate with us directly http://discord.gg/prismml
Though this says mainline llama.cpp has their patches for Metal and CPU backends, so maybe it's simply "use current llama.cpp" if you have a Mac or fast enough CPU/memor to use the CPU backend.
The fork runs fine for me. The model gets very notably stuck in a reasoning loop on one of my simple tests, though it might be that it has the same issues with setting reasoning effort high.
On my M1 Max I still think the MoE Qwen 3.6 and Gemma 4 models are the best options. And I am far from convinced that the 35B is actually worse; it gets stuck in reasoning loops much less often than 27B in my experience.
Binary: 9 t/s prompt, 6 t/s generation. Ternary: 0.8 t/s prompt, 0.7 t/s generation. It looks like CPU inference for ternary isn't optimized yet.
[1]: https://huggingface.co/models?sort=trending&search=fastconte...
The system as a whole is meant to support that use case, where each task (ticket in its jargon) can be tackled using a custom workflow that can each use a different agent/llm (so, it should support local LLMs if you have configured your coding agent to use them).
Sidenote: it's still not where I want, but getting there...
It's going amazingly. The orchestrator holds the big knowledge from grilling and also enables me to do more grilling to refine the specs. The trick was to check and re-align after each phase was developed. Also carefully defining each workflow step explicitly otherwise it makes mistakes like trying to self review the code. Also needed to define very explicit contracts. I chose the same surfaces as the matt skill set.
Edit: So life got in the way and I couldn't isolate it. But I have made this project public and committed the dev and code review agents here.
https://github.com/rick2047/meuseum-game-new/tree/main/.open...
The delegate builder is too ugly to share yet. But to be honest just use the normal builder with a frontier model and instruct it to delegate to developer any developer task. That works well.
Of course you have to install the skills [1]
I wish KV-cache memory usage and related optimizations were discussed more clearly in new model announcements and demos.
I assume that in addition to the significant memory savings, this should also lead to much simpler matrix multiplication operations? Could models like these run on CPU's efficiently, or does the geometry of the compute mean GPU's are still a better choice?
Matmul is very well parallelized. More, lower power cores will always pay off handsomely.
Ternary Bonsai 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight.
1-bit Bonsai 27B uses binary {−1, +1} weights with the same group-wise scaling, giving 1.125 effective bits per weight.
is it a float? if so, how many bits is the float?
I've never heard of a bit ever having more than two possible values
The way they do it is packing like the other comment says.
Each byte represents 5 trinary values instead of 8 binary, and there is a little bit of waste.
e.g. 5 trits (243 states) into a byte gives 1.6 bits per trit: https://compilade.net/blog/ternary-packing
You can beat the efficiency of 5 trits in 8 bits (1.6) with as few as 17 trits in 27 bits (~1.588), but once you account for rounding up to a whole number of bytes for practical reasons, then beating the efficiency requires going to at least 111 trits in 176 bits (~1.586), or perhaps more practically for fast unpacking, 161 trits in 256 bits (~1.59).
At that level, even if you have, say, 27B trits, the more efficient encodings would save something like 38-45MB (theoretical limit ~48MB), likely at the cost of some slowdown.
Their fork corrects the second inefficiency by using a group size of 128, but still uses 2-bit weights AFAICT.
It's possible to pack 5 trits into a byte, but the unpacking is not very efficient. Another recent idea is to add the constraint that exactly one weight in each group of four be zero, which gives exactly 32 possible states, so it fits in 5 bits.
If that's the case then why not just train at Q2? I guess the counterargument is that then you lose the nice properties of things like the FairyFuse kernels. I wish there were some good discussions of these trade off.
It's not represented by a "bit", binary digit with value of 0 or 1; but with a "trit", ternary digit with value of {−1, 0, +1}.
> 25g protein for "spaghetti, carrots, peppers, garlic and herbs"?
Maybe it assumed the pasta was some kind of protein chickpea pasta? =P it definitely seems wrong.
Even the lowest quality one has 12g/100g.
For the price, from a protein perspective, it would be cheaper to get it elsewhere.
https://shopfelicettipasta.com/products/monograno-felicetti-...
[0]: https://m.media-amazon.com/images/I/81UbkYz+pZL._SL1500_.jpg
> Main ingredients: Durum wheat flour, water/eggs
May be going out on a limb here but I think it might be the eggs?
/satire, obviously ;)
https://www.weekand.com/healthy-living/article/protein-durum...
> the load on your machine is gonna make doing other stuff while it’s running painful
Is your question about something else?
The biggest issue I've found is absent mindedly opening YouTube or the like that spike ram requirements and freezing the system up. But that's a me problem
eg asked it what Signoz is. It reckoned it is a woocommerce/shopify competitor aimed at India market
Sounds like the model is not following a proper probabilistic choice here, so maybe more a programming error than a model training error.
When I saw 27b on a phone, I thought not fitting, big phone, or aggressive quant. NVFP4 still takes 27G before KV cache.
I find these style of models are great, but fail hard, and fail randomly. I'd be hesitant to use it for a daily driver, but I'm using dual 3060s, so it's not like I'm quantizing a frontier model here.
How do you find the overall experience? And do you have any special sauce or recommendations for going this route?
The 2 bit quants are really good. I have a lot of memory so I can squeeze it all in at ~80gb.
model | disk | wikitext | gsm8k (match/error)
baseline | 55G | 8.00 | 0.50/0.09
nvfp4-gptq | 27G | 8.25 | 0.47/0.9
nvfp4a16-gptq | 27G | 8.11 | 0.53/0.9
bonsai-4bit | 19G | 16.75 | 0/0 (eval bug?)
Looks like they quant'd too hard at 4 bits, can't imagine the ternary being any good based on this. I'm also not sure what is up with the gsm8k, their benchmarks show something different, but they are using another eval tool. I'll have to add it to my setup. Also why I'm building a setup instead of taking model devs word for benchmarks. (https://github.com/modelscope/evalscope)Code if you'd like to reproduce or try other test sets: https://github.com/verdverm/quantr (lightly tuned to a single oem spark, probably possible in 32-48G)
Good paper to understand the effects of quant regimes across model families and tasks: https://arxiv.org/abs/2402.18158 (Evaluating Quantized Large Language Models - 2024 ICML)
Software was already complex, now we are adding highly non deterministic elements into the mix
At this point all the different quantization and 'compression' (look at MPO applied to LLMs...) techniques start feeling a bit like snake oil. It's just gut feeling - or scores on benchmarks models are optimized for - what ends up deciding whether a technique is good enough or not.
However, this is the 5th product post in 2 weeks that proclaims that AI use is shifting, and why [insert tradeoffs] are the perfect fit.
Paradigms shift don't happen in the release announcements.
I suspect this is an AI-ism, making all the release posts sound so paradigmshiftery.
Likely apples / oranges
OS: WSL2 on Windows 10
I'm interested in the CPU inference application of these models with things like the FairyFuse kernels.
I've tried trillim previously but was disappointed that i got higher tok/s just with similar sized models through ollama using just Q4_K_M quants.
I see there is bitnet.cpp and litespark-inference. What else should i look at?
The LLM style of writing is just very distracting to read. “It unlocks X”, “Y changes the equation”, and why is there always something shifting? Makes my eyes glaze over in an otherwise interesting post.
I was using vscode and it seemed to interperet the system prompt right and then started actually inspecting and doing stuff.
Unfortunately the vscode system prompt is 24000 tokens, and I was getting 100 at beginning, 69 by the end of it, but honestly I'm super impressed. Great work team 1
Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better at the size-performance pareto and provide people with strictly more possibilities.
It always auto disconnects.
My brief experiments with the ternary version suggest that it broadly meets their claim to be a 27B model that fits in much less RAM, that is for sure. It is about as fast as the underlying Qwen 27B but it gets stuck in reasoning loops quite easily.
Do they have plans to bring even bigger models down to ~16GB VRAM so that more consumer hardware might be useful?
— Hey, model, see this fake-ass stock photo of a variety of spices, vegetables, and spaghetti? What meal can I make with this?
— Just cook everything.
— I’m a complete noob. I can’t even fathom how to cook those things. Help me!
— Sure sure. First boil the spaghetti completely and drain. Only after that, while it’s getting cold, you need to sauté (good luck knowing what that is if you don’t even know how to cook spaghetti) the garlic and carrots at this specific temperature (good luck figuring out how to do that on a stove). Despite having mentioned the peppers and herbs in the previous message, I’m not going to tell you what to do with those. Just chew them raw or something, I guess.
The demo shows that the model can answer, but the answers are frankly bad. Here’s what you could’ve done instead faster with better results: a web search for “spaghetti carrots peppers”. Don’t even need to add “recipe”.
Presumably you’ve been using the model as you develop it, why not show something real and useful instead of a generic, unrealistic and uninteresting scenario that above all makes it look incompetent? Show something that genuinely surprised you positively.
This is actually difficult because there are so many invented words in the poem which have extremely low frequencies in the training data. So a model that can do it properly is likely to be very good in other ways. Qwen-3.6-27B can do this until it gets overly quantized.
you also might single handedly pop the hyperscaler investment and capital projects! that's the whole AI bubble essentially!
I can just see their image tool on the app store
Available on HuggingFace: https://huggingface.co/collections/prism-ml/bonsai-27b
The article is about running it on a phone though, and shows an app with their branding running this in text mode on a phone. I'm asking where can I find this app to try what is being demonstrated in this article & video? Appstore only has an image gen app by them and other MLX apps I've tried don't seem to support this model
> Ornith-1.0-9B, which can be easily deployed on edge devices, matches or exceeds the performance of much larger models such as Gemma 4-31B and Qwen 3.6 35B.
https://deep-reinforce.com/ornith_1_0.html
Only tried it so much so far; it did a little better than Qwen 9B
The title says it's 27B grade running on a phone and what I was comparing it to in my mind was a model that runs at 35B grade that could presumably run on a phone "better"?
edit: I asked AI for the difference and understand a little better, thanks for the heads up to learn the difference between models... I think the thing was, although ornith was created for a specific agentic purpose, it was still outperforming a previous generalist model I had running locally (so in my mind I thought it was still a better local model) - I'd like to try bonsai out if I can figure out how to run it lol
huggingface.co/Tivaphraen/Geryon-9B-v1