This is wrong for at least two reasons:
1) The data scraped from the web is filtered, deduped, ranked, and cleaned in a variety of other ways before being used for training. While quantity is necessary for a model to learn the structure of language, quality is even more important once you have a model that can produce coherent output so a lot of work goes into grooming the data.
2) There have been a bunch of papers and training runs that show synthetic data created specifically for training a model is as good or better than scraped human produced data. The importance of scraped web content is quickly declining because you can now generate infinite higher quality examples using existing trained models. The only relevance it has now is for knowledge about new developments and that is a much easier stream to filter since most of the important stuff comes from official sources and you don't need as many variations since you can just generate your own using one or more examples.