Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA
LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4
Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6 Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA
LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4
Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6You're talking apples to oranges. The "Alpaca method" is a dataset generation method. Nothing about Alpaca's training method is novel, interesting, or efficient. Alpaca used the same standard training method everyone else uses, A100 clusters.
If you mean LoRA/PEFT training which people used to replicate Alpaca then that is also apples to oranges because LoRA/PEFT is a finetuning method not a pre-training method.
Presumably…
The Pythia models are also worth checking out, they might be better than or matched to CerebrasGPTs at each size (although they warn it is not intended for deployment).
Conclusion: the landscape of top open models remains unchanged.
> It would be interesting to know why you chose those FLOPS targets, unfortunately it looks like the models are quite under pre-trained (260B tokens for 13B model)
> We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for our customers that seek that performance and want to get LLaMA-like quality with a commercial license
Which is the point made elsewhere in these comments, e.g. https://news.ycombinator.com/item?id=35344192, and also usefully shows how open Cerebras are. They're pretty open, but not as much as they would be if they were optimising for filling in other companies' moats.