The only real difference between then and now is that OpenAI's models are significantly better than my models from 2015, and they have that because well, they can afford to pile on more data. TBH, I never even considered using a large proportion of the whole internet as a training set as even remotely possible due to the sheer mind boggling costs.
Even now, to go through about 10% of The Pile would cost me way too much money.