2) Almost, Chinchilla concerns itself with minimizing cost(training) + cost(inference). The expectation most people have is that there's an insane amount of inference compute, and training is maybe 1% of that. The point of the Chinchilla paper is that that's not true, training uses such insane amounts of compute that despite all model inference by half the internet for a year or two is a huge amount of compute, training is still a very decent percentage of that. I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model.
People expect it works like compilers, that making a compiler produce 1% faster code is worth 50 highly-paid SWEs because while an individual program run isn't exactly expensive, the time and resources spent running programs is astronomically larger than the time and resources spent developing compilers.
The thing is most of the optimizations we know don't work during training. You can't quantize, you can't MoE (well, you can, obviously, but it doesn't save any training computation. In fact it increases training cost)
At 50-50, having a 10% cheaper-to-train model justifies 10% more expensive inference.
3) Combining both arguments ... at this point people should probably realize that Facebook's LLama is really an attack on Google (which is at least partially working, elon musk is tweeting about it)
If Facebook really doesn't care about AI (or ... about as much as, say, netflix does. Not zero, but as long as they beat reddit's efforts they feel very comfortable), but Zuck does care about destroying Google, the calculus changes. Zuckerberg may not want the best possible AI, he may want as many scammers as possible trying to Game the Google search quality team, to present them with challenges faster than they can adapt. Then the training cost becomes a moot point.
Hmmm, I should send my CV to meta ...