The Anatomy of a Large-Scale Hypertextual Web Search Engine (1998) [pdf]
infolab.stanford.edu
infolab.stanford.edu
"Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For example, in our prototype search engine one of the top results for cellular phone is "The Effect of Cellular Phone Use Upon Driver Attention", a study which explains in great detail the distractions and risk associated with conversing on a cell phone while driving. This search result came up first because of its high importance as judged by the PageRank algorithm, an approximation of citation importance on the web [Page, 98]. It is clear that a search engine which was taking money for showing cellular phone ads would have difficulty justifying the page that our system returned to its paying advertisers. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers."
The challenge is competition. Or rather, the lack thereof. I worked for Google when the last credible challenger threw in the towel and search got pushed closer to being a monopoly / monoculture. At the time Microsoft hadn't gotten their act together. Even as an enthusiastic imbiber of the Google cool-aid, this was deeply worrying.
Google distinguished itself because it had to in order to compete. Making suck noticably less and being in tune with the users used to be enough. Because the competition sucked at this. Google no longer has to do that. Jeff Dean doesn't lie awake at night worrying about web search competitors. There was a time when he probably did (if I remember various accounts correctly). I doubt he loses any sleep over it now.
Lots of people have tried to do distributed search. Lots and lots. The reason you don't hear about them is that they tend to not produce convincing results. Or in fact any results. In search quality, privacy, fairness. It is much, much harder than you might be tempted to think. And it exists within a field that gets harder every single year since the problem size and complexity keeps growing as the web grows; and grows more complex.
If you believe in distributed search I think you should give it a go. Just because everyone has failed before you doesn't necessarily mean it is impossible to do. But it is worth remembering that it is a class of problem that is a lot harder than centralized search. So saying it is the solution is like saying cold fusion is the solution to all our energy needs: it might have been if we knew how to achieve it. But we don't.
https://courses.cs.washington.edu/courses/cse454/09sp/papers...
There is a updated paper describing Mercators design about a year later when it moved to scale on multiple machines and became a continuous crawler:
https://www.hpl.hp.com/techreports/Compaq-DEC/SRC-RR-173.pdf
Also helpful is the entire Modern Informational Retrieval textbook which is available online and dives into indexing, full text search, etc
https://nlp.stanford.edu/IR-book/information-retrieval-book....
It’s fun because many challenges DEC, Google, and others had in the late 90s can largely be solved by modern machines (disk seek times, having to do appends in bulk, storaging reverse indexes in RAM, etc)
Starting with something dumb and replacing it with something smarter whenever the dumb thing doesn't work does seem to go a long way.
Since it is very difficult even for experts to evaluate search
engines, search engine bias is particularly insidious. A good
example was OpenText, which was reported to be selling companies the
right to be listed at the top of the search results for particular
queries. This type of bias is much more insidious than advertising,
because it is not clear who "deserves" to be there, and who is
willing to pay money to be listed. This business model resulted in
an uproar, and OpenText has ceased to be a viable search engine.
But less blatant bias are likely to be tolerated by the market. For
example, a search engine could add a small factor to search results
from "friendly" companies, and subtract a factor from results from
competitors. This type of bias is very difficult to detect but could
still have a significant effect on the market.
Sounds like a plan! ヽ($_$)ノTo the point, search engines have become the window on knowledge for most people, and thus allow something that goes beyond that orwellian process, an almost direct control over "history", live. For many subjects, it has become almost impossible to find results that are more than 50 months old. You can spend some time trying to do research. But this process is very long, and it is only one subject.
So to me a truly open, transparent and independent indexing and search system is one of the most important technological problems to solve.
Because of the vast number of people coming on line, there are always those who do not know what a crawler is, because this is the first one they have seen. Almost daily, we receive an email something like, "Wow, you looked at a lot of pages from my web site. How did you like it?" There are also some people who do not know about the robots exclusion protocol, and think their page should be protected from indexing by a statement like, "This page is copyrighted and should not be indexed", which needless to say is difficult for web crawlers to understand. Also, because of the huge amount of data involved, unexpected things will happen. For example, our system tried to crawl an online game. This resulted in lots of garbage messages in the middle of their game!