I'm aware that CGI has been around for a long time, but it was bog slow since it forked new processes to deal with every call, it wasn't really feasible to build the sort of dynamic websites we do today until stuff like PHP entered the market.
The relative proportion of static files versus dynamic web content have probably reversed.
> This really is a case of the tech industry trying to redefine "search" in a way that gives them editorial power.
This isn't a new development at all. The signal to noise ratio problem goes back to the late '90s, and ruthlessly removing noise was arguably one of the moves that allowed Google to break through.
Unfortunately they seem to have forgotten why they became a success.
From http://infolab.stanford.edu/~backrub/google.html
> [...] In 1994, some people believed that a complete search index would make it possible to find anything easily. [...] However, the Web of 1997 is quite different. Anyone who has used a search engine recently, can readily testify that the completeness of the index is not the only factor in the quality of search results.
--
It's difficult to conceptualize these numbers, but we can use English Wikipedia as a measuring stick. It has of the order ten million pages, and covers most topics you would ever want to read about. Maybe there are some niche things here and there that are missing, but overall, it's a relatively complete coverage of human ideas and interests. How many English Wikipedia Units do you need, for your search engine to feel complete, assuming 100% signal and 0% noise? Do you need one? Ten? A hundred?
A "big search engine" crawl is over a hundred thousand English Wikipedia Units. How much of that is interesting? How much has been written by humans, has ever been looked at by humans?