Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.
2,676 karma · joined February 8, 2017
Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.
Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is one, it seems like it would be useful to track.
“The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”
How does soup auto tune the hyper parameters and make some of these more complex training decisions?
What the hell happened to stack overflow? I don’t think I’ve gotten a google result for stack overflow in nearly 2 years. How have they just forgot how to make competent search?
I didn’t think it’d happen so quickly but I honestly get better results from DuckDuckGo at this point.
I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maartengrootendorst.com/p/a-visual-guide-...
In reference to the difficulties with applying this to autoregressive LLMs - I wonder if these type of hybrids might be a good candidate for this approach.
"Neutrino-1 8B was trained natively in its shipping format. There is no full-precision product model that was rounded afterward: the ternary representation is the medium the weights learned in, and the training methods that hold this quality at this depth are the lab’s unpublished work. The findings below are the part that travels."
This statement seems misleading at best.
Both the model page and the release page are basically unintelligible - I don't have a ton of faith in the work here, at least PrismML write coherent releases for their models.
Edit: Another beautiful piece of prose here, I almost wonder if they used the 8b model to generate the content for this release...
"Across the 6.95B coded weights, 62.63% sit at zero and the remainder splits 18.68% plus to 18.69% minus: sign-balanced to a hundredth of a point with no constraint asking for it."
PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsai
I wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764
From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."
The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...
I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.
It is a great solution for places where heating is the primary use of electricity and I hope it finds broader applications.
Anyhow, this kinda reminds me of that quote about architecture: "We replaced our monolith with micro services so that every outage could be more like a murder mystery."
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
This is pretty impressive.