IsoFLOP curves of large language models are flat
severelytheoretical.wordpress.com
severelytheoretical.wordpress.com
Of course, ultimately we will figure out scaling laws for LLMs trained on multiple epochs of data, but not today.
By duplication I mean if context length is N there is many sequence of N word that are not unique.
How is this loss calculated though? Since it is called "loss" and not "performance metric", I'm going to assume it is the teacher forced cross entropy loss.
I'm not too familiar with LLM training but having been doing some fair amount of seq2seq training lately in other domains I've observed that the relationship between "loss" and autoregressive inference performance gets very narrow towards the end of training. What I mean is that smaller and smaller reductions in loss lead to better improvements in the autoregressive output. So at least I have some doubt that in practice that 1% loss improvement is not actually incredibly significant with respect to how well the model actually performs at inference time.
But I'm pretty interested in this topic and if people here have observed this or the contrary I'd be curious to know.
Here is a quick screenshot if you are lazy.
https://snipboard.io/C6mipQ.jpg
And here is the paper if you want to dig deep into the highest quality published research on LLM frontier.
https://ai.meta.com/research/publications/the-llama-3-herd-o...