If 13B is good... I wonder if this will catch on in the finetuning community.
People care less about the LLaMA license than you'd think, and this is also about the time new models with "improved" architectures (like Falcon) should start popping up.
If 13B is good... I wonder if this will catch on in the finetuning community.
People care less about the LLaMA license than you'd think, and this is also about the time new models with "improved" architectures (like Falcon) should start popping up.
Weights being subject to copyright was never tested in court. Businesses have wisely steered away from using leaked weights for commercial purposes. With OpenLLaMA we can expect novel products that incorporate these small but capable LLMs, for example as game NPCs.
My knowledge of neural nets and AI is just lacking.
https://user-images.githubusercontent.com/48489457/243093269...
You can see that "Perplexity" goes down as Model Size goes up.
Their point was a lot of research that had been showing jumps in performance at certain "breakpoints" for lack of a better word were the result of badly selected metrics versus a case of suddenly emergent behaviour.
The nice thing about that research is it suggests that if you are able to try something on a smaller model it will scale nicely to a bigger model.
The key insight from the abstract, "Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance."
Yup that is another recent and interesting development for sure!
Also note that the plots in the appendix contain some obvious errors, so you definitely want to wait for a peer reviewed version of this paper (if it ever survives review).
Their point was a lot of papers use those weird metrics and it contributes to the appearance of emergent ability, when in reality its just the bad metrics.
Nothing you've said so far disagrees with either my understanding or the conclusion of the paper I linked.
Several factors influence performance beyond parameter count, notable ones include: training corpus quality, training flops, and the downstream task.
It depends on how much compute you are spending on training and how big of a model you’re talking about.
There’s a “minimum” tokens/parameter for increasing size to be effective at improving loss/perplexity, so as you go up in parameters you generally will have to broaden your corpus which may lower it’s quality (e.g. tweets/reddit posts vs books/articles).
This effect isn’t as significant at 65B parameters as there is still enough high quality training data but if you’re talking 1T the corpus will (probably, I haven’t tried this/seen this done) by necessity be significantly poorer quality as you will overfit by simply repeating (it’s only beneficial so many times to repeat).
As a general rule, when validation loss/perplexity are the same in two models of different sizes downstream performance seems to be also generally the same (this was briefly explored in the PaLM 2 paper by Google) although it doesn’t correlate perfectly.
Practically speaking, this translates into a bigger model is better for applications we’re generally talking about. It just may not hold infinitely which we’re starting to see evidence of.
It’s definitely not linear though, you can look at some of the OpenLLaMA benchmarks (without getting into the weeds of if current benchmarks are representative) and the accuracy improvements in even 13B is not that significant (noting here all models were trained for the smear number of tokens, so relatively overtraining the smaller models).
There are probably some threshold parameter sizes that make a big difference but it’s still being determined.
But note thst Meta's LLaMA 33B and 65B were trained with more tokens (1.2 trillion?) than the 13B and 7B models (1 trillion).
And subjectively, the larger parameter models do indeed feel "smarter" beyond what objective metrics would suggest.
https://huggingface.co/models?sort=downloads&search=13B
It is a mess, but its pretty much the only destination for finetunes from various communities/institutions.