Bing versus Google, some observations
jacquesmattheij.com
jacquesmattheij.com
Google claim that they saw lots of (less obvious) evidence of Bing mining search results from Google before they began their sneaky test, and the point of the test was simply to confirm that Microsoft is doing what they thought.
This is not about whether Bing is easy to "game" -- whether Google can get nonsense into Bing's index by sneaky means. It's about whether real Bing searches commonly derive their results from Google.
Imagine that I think you're reading my email and using the information in it to play the stock market (maybe I have secret insider information about some companies, or something). So I do a test: I arrange to be sent a bunch of email that, if you acted on it, would make you buy particular companies' shares that you'd otherwise have no reason even to have heard of. And, lo, you do that for 10% of the companies involved. Would anyone, looking at that, say that the real news is that I was unable to "game" your stock market transactions effectively, and that your spam filters caught 90% of the junk I tried to inject into your information?
BTW, that's not necessarily about whether they are "lying". The other evidence they think they have could be convincing to them (because, e.g., they already have a strong predisposition to believe that they are inherently far better than anyone else, and therefore anyone building a competitive search engine could only possibly be copying them - this is a caricature, but you get the idea) but not necessarily convincing to others.
Yes, it could be a side effect of some machine learning algorithm. But Bing never explained that, which would have been very easy if it was the case.
What Google presented was enough to show that there exist possible situations where an individual (query, URL) pair's presence in Google's search results causes it to appear in Bing's. That definitely demonstrates that Google has a nonzero influence on Bing, and you can call it "copying" if you like - on the level of individual (query, URL) pairs.
I personally don't think there's anything wrong with this, in itself. I think an individual (query, URL) pair is small enough to be 'fair use', more or less. But it's another matter (to me) if this sort of thing is happening often enough, and in important enough cases, to have a strong aggregate effect on the Bing search engine as a whole. Google has insinuated that they believe this to be the case, but they haven't shown it. That was what the post above mine was about and that was what mine was about.
If you think that what they actually did demonstrate is bad enough in itself, none of this matters.
Suppose your financial advisor (Google) was suspecting that someone (Bing) was stealing their confidential financial reports on stocks. Suppose your financial advisor told you (Google Engineer) to buy 100 random shares (search and click on 100 specific search terms) and see if the suspect (Bing) acted on it.
Even if the suspect bought 100% of the shares (Bing indexes all the search terms with the irrelevant links), you still haven't proven the suspect is stealing information from the financial advisor because there's more than one source this information could have come from. It could have come from you (Google Engineer doing the clicking) or it could have come from your financial advisor (Google itself). A way to solve this issue is if you (Google engineer) had another financial advisor (another website) which told you to buy certain companies. If the suspect didn't act on those shares then you would have MUCH more conclusive evidence that the suspect was stealing from the financial advisor represented by Google.
(Determining the degree to which Google's methods are similar to a 419 scam is an exercise for the reader).
I see a difference between aggregating content and presenting it and mentioning the source and just plain copying (such as spell corrections) with no mention of no source.
But I think the author focused too much in the "9%" part. Who knows what Google did with the other 91%? Maybe they were trying different approaches, which actually would be the most sensible thing to do.
That is sensible, but can easily slide into an ends justifying the means. In other words, we know that the engineers were tasked with proving that Microsoft was copying Google. That's a different task than figuring out what the Bing toolbar does with data from the Google search page.
Given the low success rate, it is not implausible that the engineers pushed whatever ground rules there were (if any) in the pursuit of evidence. It seems to me that is the most plausible explanation for the uncertainty about whether it was 7, 8, or 9 cases.
To go beyond what is easy for the media to report, it is reasonable to expect that a company in Google's position has twenty or more full time engineers analyzing their competitor's products. The story about "torsoraphy" really only makes sense if Google has such a program. Seriously, is anyone surprised that Google and Bing analyze each other's engineering?
[WildSpeculation] One or more Google engineers tasked with analyzing the competition discovered "torsoraphy" connection and identified its correlation with the Bing toolbar - however, keep in mind that Google has not claimed that the "torsoraphy" naturally occurred in the wild. [/WildSpeculation]
[GoogleClaim] Based on the "torsoraphy" discovery, one or more Google engineers hard coded web pages - presumably with permission from senior managers since a leak that it was done casually has such serious blowback potential - twenty Google engineers armed with laptops were tasked with creating top keyword rankings.The project was at least active for two weeks over the traditional Christmas Holiday[/GoogleClaim]
[WildSpeculation] The Google engineers, surprised at their initial lack of success tried increasingly diverse and aggressive methods as time passed despite the diversity of IP's they used as they traveled over the Christmas Holidays. Being good hackers some even tried exotic techniques perhaps even creating manipulating existing web pages to influence the search rankings. On advice of lawyers, Google is not comfortable accusing Microsoft of copying in these cases.
A month later, Google plans to announce the Android Marketplace while Apple is going to proclaim that Rupert Murdoch is the future of journalism. Google is going head to head with the reigning PR champion of the world. They turn the experiment into a torrid story and release it on February 1.
Larry Page knows the difference between a Founder from Stanford and a former IBM'er from a cow college in Alabama. It's no match. Binggate drowns out the traditional eve-of-event Apple adulation. On the day of the events, the mudslinging is far sexier than "and it has a hundred pages" for the tech press.[/WildSpeculation]
How bad was it yesterday for Apple? Stories about Apple's triumph with The Daily didn't make the front page of HN. The tech press even found Microsoft more interesting than Apple yesterday. Binggate was Googles attemt to kill two birds with one stone.
So that 'excellent point' only shows that Google even promotes products of its competitor. I don't think it says anything about Google "ignoring" copyright. Heck, you could even look into that book at Amazon for free: http://www.amazon.com/gp/reader/0735622841
Along your reasoning, every library or friend that shows you a book is ignoring copyright.
Microsoft "copies" the results. Does not say where they came from and presents them as their own. In this context, that would be copying all the text from that book, removing all the branding and creating a new book saying that Google wrote it.
I have no issue whatsoever with Bing analyzing the data sent by opt-in users that is further processed to help increase its search relevancy results.
I would not care if Bing even sent Google synthetic and organic search terms directly, to analyze and make use of the results.
Remember, Google uses you (the searcher) as a product that is sold to the real customers: the adwords and adsense clients.
Google scans the net of content that they do not own, and makes an enormous amount of profit in return.
Google owns nothing. You owe Google nothing.
But if "Bing even sent Google synthetic and organic search terms, to analyze and make use of the results", I would suggest they add 'powered by Google' to their search result page.
Further, I don't think it makes a difference if Bing is querying Google directly or using opt-in users to do so. In any case, their copying search results efforts from Google.
Again, I don't see why it's an issue for Bing to analyze and make use of the behavior of Bing toolbar users that opted in.
So what if their Google search behavior holds good weight as a metric?
I've always felt that, yes, Google collects a lot of personal data, but they're up front about the collection and they give me good value in return - which is sadly not my general expectation of data collection.
On the other hand, do we know the percentage of searches that actually get redirects as results? The 'honeypot' was rather small, so the redirects just might have been too little to appear as a signal in Bing.
The lack of google redirects in bing's results doesn't look like proof, or even a smoking gun to me.
(Answering for Jacques, since he said he won't be here anymore to answer for himself: http://jacquesmattheij.com/Tell+HN%3A+So+Long+and+Thanks+for... )