Solar-10B
twitter.com
twitter.com
I have tried a lot of these LLMs locally and I have rarely ever found the smaller models to outperform the bigger models when used casually/generically. The same is generally true for 1.5 vs SDXL the moment the prompt strays away from some variant of beautiful woman.
On the LLM side, there's a similar phenomenon where a huge amount of training could be on SEO spam tags, or /r/counting or producing nonsense or generating obfuscated C or reciting excerpts of Shakespeare word for word, etc.
If what you're measuring is generally more useful then you can end up with a better model using these methods, likely at the cost worse performance on things that you aren't measuring.
I would gladly take a 10b parameter coding assistant that doesn't know how to write in iambic pentameter or recite digits of pi, or translate words from Swahili to Turkish etc. but is much better at code completion.
If a model forgets how to speak French but gets much better at generating unit tests, that might be perfectly fine for the type of work we want the model to perform.
The problem is we can't easily know what the model "forgets" when it gets better at doing something else. The best thing we can do is benchmark/measure their output and hope that those benchmarks cover what users care about.
I suspect high quality benchmarks will quickly become almost as important as the tuning process itself.
But the top models right now are almost all under 70B. Most are 7B, and the top is 10B. If the benchmarks are even remotely accurate then this is rather wild.
Apparently multiple groups found different "secret sauces", names upstage and whatever UNA is?
I can’t imagine this is anything but selection bias.
Finetunes rarely led to "Top 5 performance" for the small ones. Previously the top 10+ were all 70B, with maybe a few 30B in there. There were nearly no 13B's, let alone 7B.
The Zephyr-7b-β was one of the best 7B mistral 0.1 finetunes the past month and a half, and that didn't beat most 70B's.
Even at 7B there are few foundational models as even those take a relatively large amount of money. The only decent one for months has been 7B mistral which again didn't come that close to 70B performance.