The Barbell Effect of Machine Learning
medium.com
medium.com
This seem like nonsense. I'm I wrong in thinking that when Google still used PageRank that it wasn't based on ML?
Generally how much data you need for ML varies with the problem and to say that it's all in the data and large server farms is perhaps a nice talking point for VCs but can be completely true, completely false, or anywhere in-between depending on the problem.
PageRank meant the system was less easily gamed and produced higher quality results. The real magic was combining that with their revolutionary data center operations and being able to treat a whole lab of machines as a single device which ran via map/reduce.
The clean design helped as well but it was incremental. Google's advantage was that they produced better results faster with lower operating costs. Better and faster were good for customers. Then the rollout of relevant ads meant that Google could turn a profit even at lower cpm than competitors and their ads didn't have to be intrusive or distracting. The combination put them on much surer footing than their competitors so they kept growing, increasing revenue and profit while their competitors struggled.
The replication recipe - be really clever and try really hard.
The interesting thing about a lot of the big successful tech companies is that they blend things that are often difficult to find together naturally, which makes their success difficult to replicate. Google basically runs on alien technology, they're sort of the ... I guess Tesla or SpaceX of web-apps. It's possible to do what they do, but it's so outside the range of the traditional way of doing things that it would take a tremendous amount of talent, vision, and determination to pull it off. Though google is also very dysfunctional in places so from a business sense it's not actually as hard to compete against them as it may seem. Apple is built on a weird marriage of solid aesthetic sensibilities and solid engineering fundamentals. And then Amazon which marries cutting edge fulfillment to a solid web store front-end.
For a lot of these things it'd be straightforward to try to build one aspect, but getting all of them together is a huge challenge.
However, it is learnt from the data: links from one document to another are recorded in a huge adjacency matrix (sort of) and the PageRanks are dominant eigenvector of that matrix.
The original paper never mentions AOL or Yahoo (except as an example of a popular website), but does describe how they built their own crawler (see here: http://ilpubs.stanford.edu:8090/422/1/1999-66.pdf) but I suppose it's possible they got data from those companies later.
But the traffic Google got from the Yahoo and AOL deals did allow them to see user behavior, and that did allow them to improve their search ranking algorithms (which of course wasn't really PageRank anymore).
Though as others have said, after the start gun went off, a lot of other factors came into play because PageRank, like other algorithms, is easily gamed on it's own.
Edit: after thinking about this a bit more, I guess you could in fact think of e.g. k-means clustering as just a very advanced form of descriptive statistics, perhaps not fundamentally different from calculating a mean or a kernel density estimate. And in that sense PageRank would be unsupervised learning too, but it still feels to me like that's obscuring rather than clarifying how it works?
That said, one interpretation of the PageRank model is of a user randomly clicking from page to page, occasionally giving up and starting over. AOL would have access to data like this, which might help them refine the "damping" term in the PageRank model.
On the other hand, what kind of data could be so secret that only the big companies could accumulate? If you have access to a service, you can mine it and construct a dataset. Then you can train your own model to imitate their results. If this process is done right, it becomes easy to do transfer learning from other AI systems, even if they keep their algorithms secret.
For example, Google's latest POS tagger SyntaxNet, which made a splash a few days ago, was trained on the results of the Stanford parser. Interestingly, the student model surpassed the teacher model - probably because it was better at cancelling out errors in the training set and generalizing better from examples.
If data is the end game of these companies, then it will be a target of espionage, leaks and public disclosures. A hard drive containing the prized dataset of some company can be copied in a few minutes and if one copy escapes, then it is circulated and becomes public domain. So it's hard to protect data. Data likes to flow (remember the Ok Cupid dataset - that was crawled without permission from the behind the login wall?). Flowing data can be turned back into machine learning models.
That's not going to stop some hacker could crawl OkCupid and balckmail some people. But if Google or Microsoft suspect FooStartup Inc. is using one of their datasets in its service, then the lawyers will descend.
This is flat wrong. No algo, no intelligence or synthetic congnition
> The dramatic rise of Google provides a glimpse into what this kind of privileged access can enable. What allowed Google to rapidly take over the search market was not primarily its PageRank algorithm or clean interface, but these factors in combination with its early access to the data sets of AOL and Yahoo, which enabled it to train PageRank on the best available data on the planet and become substantially better at determining search relevance than any other product.
This is so wrong on so many fronts. A) They has open access to the web just like DMOZ and Yahoo and many others via crawling systems. B) They were attractive to software engineers who in turn made their IT depts in large corps switch to Google as the default search engine C) They stole the ad matching algo from Bill Gross which in turn made them successful. E) too many other factors to list
Lets also not forget that PageRank was a variant of an early link analysis algo that had been around a while.
On the second statement, I believe you both missed the point. Google was able to acquire click-through data which helped train it's rankin algorithms better by establishing the partnerships with Yahoo and AOL. It was the implicit feedback, not the crawled data that helped. But that obviously wasn't the only factor--having very smart engineers/scientists matters. But I think now that a lot more is open sourced, computing power is cheaper, etc the OPs thesis is that data is now the hardest thing to come by, and on this I fully agree.