True. I shouldn't have used a universal qualifier. I should have, "all the data possible (that one corporation can get it's hands on)" or something qualified.
The CEO and CTO of OpenAI have both said that they currently have more than 10x data than they used to train GPT-4, agreements to collect 30x more, and that collecting 100x more would not be a problem.
Do you have a source link for this?
Probably not even that. Remember that the constraints also include cost and time so it’s unlikely they just threw everything at it willy nilly.
Another avenue is training on generated text. This is likely to be important in teaching these things reasoning skills. You identify a set of reasoning tasks you want the system to learn, auto-generate hundreds of millions of texts that conform to that reasoning structure but with varying ‘objects’ of reasoning, then train the LLM on it and hope it generalises the reasoning principles. This is already proving fruitful.