However, I’m not even sure enough text data exists in the world to saturate 100T parameters. Maybe if you generated massive quantities of text with GPT-4 and used that dataset as your pre-training data. Training on the entirety of the internet then becomes just another fine tuning step. The bulk of the training could be on some 400TB dataset of generated text.