I've heard that Google and Baidu essentially started at the same time, with the same algorithm discovery (PageRank). Maybe someone can comment on if there was idea sharing or if both teams derived it independently.
I've heard that Google and Baidu essentially started at the same time, with the same algorithm discovery (PageRank). Maybe someone can comment on if there was idea sharing or if both teams derived it independently.
What I find very interesting about PageRank is how you can trade accuracy for performance. The traditional way of calculating PageRank by means of squaring a matrix iteratively until it reaches convergence gives you correct results but is sloooooow. For a modestly sized graph it could take days. But if accuracy isn't that important you can use Monte Carlo simulation and get most of the PageRank correct in a fraction of the time of the iterative method. It's also easy to parallelize.
Jon M. Kleinberg, "Authoritative sources in a hyperlinked environment," 1998, Proc. Of the 9th Annual ACM-SIAM Symposium on Discrete Algorithms, pp. 668-677.
> In 1996, while at IDD, Li created the Rankdex site-scoring algorithm for search engine page ranking, which was awarded a U.S. patent. It was the first search engine that used hyperlinks to measure the quality of websites it was indexing, predating the very similar algorithm patent filed by Google two years later in 1998.
The idea came first up in the 70's. https://www.sciencedirect.com/science/article/abs/pii/030645... and several times afterward before PageRank was developed.
That being said, Page Rank is a more a stellar example of adapting an academic idea into practice, than a statistical idea in and of itself.
Afterall, it is 'merely' the stationary distribution for a random walk over an undirected graph. I say 'merely' with a lot of respect, because the best ideas often feel simple in hindsight. But, it is that simplicity that makes them even more impressive.
I don't think this means much. The history of science and technology is full of examples of results named after someone other than the first person to find them.
https://en.wikipedia.org/wiki/List_of_examples_of_Stigler%27....
In fact, based on the other comments in this thread, it seems that Pagerank being named after Larry Page is itself one of these examples.
Implementing it and catching edge cases isn’t trivial
(I do not want to link directly to the pdf shown in the search result). Section 2.1 deals with related work: "There has been a great deal of work on academic citation analysis [Gar95]. Go man [Gof71] has published an interesting theory of how information flow in a scienti c community is an epidemic process......" (and more)
I think that paper is worth a read.