I was under the impression that the size and quality of the training dataset had a much bigger impact on performance versus the sophistication of the model, but I could be mistaken.
I was under the impression that the size and quality of the training dataset had a much bigger impact on performance versus the sophistication of the model, but I could be mistaken.
Compute is then proportional to the product of model size & data quantity.
That said, quality of data also matters a lot - OpenAI has had human labelers produce the data for their Reinforcement Learning from Human Feedback (RLHF), which has probably had a disproportionate impact on the success of ChatGPT compared to previous models, but that data is probably O(1%) of what they trained on.
At this point I'm guessing OpenAI are limited by both data & compute. Rumor has it they're training the "next big thing" now and it won't finish until December. If they had more compute they could presumably finish sooner, and if they had more data they would presumably let it train longer.
But a lot of the advancements we're seeing right now are the result of more sophisticated models [1], and one person is doing some interesting work [2] around achieving transformer-level performance with other architectures.
So it's not completely settled if more data is the answer. But it has a significant impact.
[0] https://ai.facebook.com/blog/large-language-model-llama-meta...
[1] https://en.wikipedia.org/wiki/Transformer_(machine_learning_...
Babies do something completely different. They can't walk when born. Their model is to blunder about and work things out, building the model up thro a can I do this - can I do that - why not etc. Its only through this doing learning happens.
We have calf ai right now..you ask the calf what do you want to learn next or what are you curious about and you get to see how dumb it is.