Estimating PaLM's Training Cost
blog.heim.xyz
blog.heim.xyz
In the long run this type of research will more than pay itself back in benefits to Google. NLP underpins everything they do. It would be interesting to see how OpenAI API (GPT-3) is doing in terms of revenue. They're going for the more direct method of seeking value from a trained large language model.
> your hyperparameters are bad and you should feel bad
It's amazing to me that such a big goof was missed by so many for so long. All these multimillion dollar language models and people just took the scaling laws at face value.
Cyclic learning rates work well elsewhere, but I can't think of any other case where switching made such a difference. I was completely shocked to read Chinchilla, and I'm still a little baffled - a cosine schedule? Really? That's it? You guys didn't make any other changes?
How many tokens of training data have humans produced across the entire internet, all our written works, etc? Is there such a thing as a 216 trillion token set?
[0] https://arxiv.org/abs/2203.15556 [1] https://arxiv.org/abs/2204.02311
But the real question you should be asking is, where would you get the compute to train a model that needs 216t tokens?
As long as the model size scaling improves performance alone, it makes sense to scale. Only once performance saturates, given same data, it's time to switch attention to training on as much data as possible - and hopefully such model turns out to be superhuman in writing code - improving itself on the algorithmic side, starting the singularity.
As an additional point, an AGI in this approach may well need to be trained on images and videos (both containing text in addition to other information) - which requires way more compute power. This would make enormous 'wasteful' text models a way to reach a point where enormous future models and training architectures can be reused for visual and sound data with small changes.
As a separate point, this suggests to me that 100x larger models are within the current reach of Google and other megacorps.
You know when on TNG they ask Computer for a bunch of related things and it somehow knows what they're talking about.
You’re right. Google had to: 1. Design, tape-out, manufacture, and deploy TPUv4 2. Run PaLM