HNHacker News
TopNewBestAskShowJobs

tfburns

6 karma · joined June 14, 2024

submissionscomments
tfburns··on Kolibri: A Sovereign Open-Weight Model
Tom here. Thanks for the note! I think we also interacted via X on this topic? Certainly interested to see what the comparison looks like.
tfburns··on Kolibri: A Sovereign Open-Weight Model
That's the beauty of Apache 2.0! People can just serve it :) there were a few kind community members offering this for free the last few days. I think there are still some, e.g., Tesseracted.
tfburns··on Kolibri: A Sovereign Open-Weight Model
Hi there! I'm Tom. I worked on Kolibri at Aleph Alpha :)

I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.

Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.

So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.

tfburns··on Kolibri: A Sovereign Open-Weight Model
Do models from other parts of the world underthink? :P
tfburns··on Kolibri: A Sovereign Open-Weight Model
[flagged]