Providing not just any a baseline, but a correct and useful one, is ever more important the less the model is grounded in world knowledge – misunderstandings probably compound faster if there is no general grasp of (broadly) “life on earth”, or computers, or whatever.
And secondly, I think (consumer-oriented) search becoming worse and worse is a challenge that’s mostly solvable (but far from solved!) for the big labs: (Mostly) trusted or even editorialized/reviewed sources like published work, Wikipedia, etc. is something they could index internally, it doesn’t need to come from a random blog site on the public internet. Furthermore, there’s a whole slew of companies specializing in crawling-for-LLM (i.e., bypassing bot protections) now as well.
We don't really have that much more good information now than 30 years ago. We had an explosion of garbage that google saw as an oppotunity to sell us on, rather than filter out.
Google saw it as their job to keep us on search longer, so the SEO garbage filling the search feed benefited them.
For a lot of use cases, you don't need a general purpose search engine – a search over a curated knowledge base works even better.
There are plenty of freely available data sets you can use, depending on the application; plus in many cases you will want to use internal-only knowledge bases containing non-public information (e.g. documentation for a corporation's internal systems and procedures). There are also many paid subscription domain-specific knowledge services available.