Also, these models are trained on trillions of tokens, so I’m not sure an 8B model even can overfit.
Also, these models are trained on trillions of tokens, so I’m not sure an 8B model even can overfit.
2) Almost, Chinchilla concerns itself with minimizing cost(training) + cost(inference). The expectation most people have is that there's an insane amount of inference compute, and training is maybe 1% of that. The point of the Chinchilla paper is that that's not true, training uses such insane amounts of compute that despite all model inference by half the internet for a year or two is a huge amount of compute, training is still a very decent percentage of that. I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model.
People expect it works like compilers, that making a compiler produce 1% faster code is worth 50 highly-paid SWEs because while an individual program run isn't exactly expensive, the time and resources spent running programs is astronomically larger than the time and resources spent developing compilers.
The thing is most of the optimizations we know don't work during training. You can't quantize, you can't MoE (well, you can, obviously, but it doesn't save any training computation. In fact it increases training cost)
At 50-50, having a 10% cheaper-to-train model justifies 10% more expensive inference.
3) Combining both arguments ... at this point people should probably realize that Facebook's LLama is really an attack on Google (which is at least partially working, elon musk is tweeting about it)
If Facebook really doesn't care about AI (or ... about as much as, say, netflix does. Not zero, but as long as they beat reddit's efforts they feel very comfortable), but Zuck does care about destroying Google, the calculus changes. Zuckerberg may not want the best possible AI, he may want as many scammers as possible trying to Game the Google search quality team, to present them with challenges faster than they can adapt. Then the training cost becomes a moot point.
Hmmm, I should send my CV to meta ...
Where are you getting that from? As far as I can tell, the Chinchilla paper is purely about getting the highest quality from a fixed training budget. Inference is only mentioned a couple of times in passing as a side effect of smaller models, not as the goal nor as an input to the formula. (And just to be clear: the Chinchilla paper was arguing for smaller models trained for longer, while you seem to be saying that they were arguing for larger models since the inference cost is insignificant.)
> I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model.
I do not see any such example in the paper
Chinchilla optimal scaling is not useful if you want to use the model, just if you want to beat some other model on some metric for the minimal training costs.