Pretty much captures it. That was the founding principle of Blekko which was that humans could curate a core set of 'good' websites for a topic and all the crap would not have a foot hold to show up on the page.
So to share some of the challenges with that (if anyone out there wants to try again) they are as follows:
* 'Search' means different things to different people, and we've been trained that a 'search engine' finds anything on the web (for the most part). The product Blekko built could more accurately be called a 'reference' engine which was used successfully by people trying to find facts or data and were not generally trying to find things to buy.
* Have your own advertising system, all in house, where you don't have to "revenue share" with anyone if you don't want to. At its peak Blekko was serving over 10M queries per day which, if we had owned all the advertising revenue on those searches would have kept us going and growing. That said, building an advertising system is both difficult and fraught with patent risk / bad-actor risk.
* Don't let anonymous users use the service. This is perhaps the hardest thing, most people won't give up an email address for even the most useful of services, however since you're spending money serving up search queries you don't want to waste that money serving up queries to bots and other bad actors. At any given time when I went through the logs there were between 5% up to nearly 18% of the queries were 'suspicious' or likely bots. That is 18% of your capacity you can't give to "real" humans if you can't control that traffic.
* Build a relevancy ranking system rather than a popularity ranking system. Search results have two metrics of interest, precision and recall. For a reference engine you want to focus on precision over recall. And while existing search engines use the "virtuous cycle" of search & click to track popularity (which can be an indicator of precision but is better at indicating click-baityness) build your ranking engine using NLP based evaluations.
* Your document index size should target 5 billion documents with a goal of 10 billion documents. Scale your cluster and algorithms to process a query to that index in 100mS or less.
Do that and win :-) Or find the next barrier to creating a useful way to search the web for information.