There is no victim here. They are not taking your 1st result and copying it. They are taking the result the user clicked. Obviously you didn't predict that with your algorithm or you'd have always made that the 1st result. Instead, what they're tracking is user behavior, not your raw ranking.
Obviously users give Google implicit permission to track their behavior by using your product. And similarly, by installing the Bing toolbar, they're giving Bing that permission.
This is beneath you Matt and it's beneath Google.
I think Google has a right to complain. Microsoft has resorted to these less than innovative tactics to monopolize themselves for a long time now, and it isn't fair to companies like Google who have worked their butts off (and gave 1.8 million shares - $336M in 2005 - to Stanford for the PageRank algorithm) to develop their superior product.
Perspective changes things here, which means no one is "right" or "wrong".
It's more like Linus say you can't use the code, but Google use them anyway.
Btw, Google contributed a lot to open source projects.
By installing the Bing Toolbar, users are giving permission to track their clicks. If Bing's server farm is searching Google and parsing the results then it is more like your example.
Let's put it this way... if Google hadn't bought the PageRank algorithm from Stanford and put years of work into perfecting their search results, Microsoft wouldn't have any way to track which Google search results users click. It's an unfair tactic that clearly demonstrates Microsoft's sketchiness and desire to monopolize themselves (by any means necessary, "evil" or not) wherever there's a computer.
As for the fanboy comment... I'm certainly not a fanboy but I'll let the following speak for itself: Microsoft Internet Explorer vs Google Chrome
From a search engine user's point of view, I believe this whole fiasco is ridiculous. First, it's ridiculous because Google is handling this situation very immaturely. Matt Cutts should not have confronted the VP of Bing in a way he did. Second, if I were the user of the Bing Toolbar, I gave permission to the Bing Toolbar to use my behaviors to polish my search results. I have no problem with that. Lastly, the experiments they did has more to do with "guessing what user wanted" than "what PageRank does".
I've used Bing fairly often past 6 months because of too many spams Google search results were giving back. Now that Google has fixed (or still working on) the spam problem, I'm starting to use Google again. However, what I noticed from the past 6 months is that Google search isn't so much better than Bing. This Bing Toolbar fiasco only applies to synthetic queries that I would never make.
Is Bing cheating? I don't think so. To me, they are just using another signal from user's permission. However, the definition of cheating will be different for everyone else.
We really don't know how much value Bing puts on clicks made on Google. Perhaps a lot?
What's the Google's official stand on this?
If the case described above were true, then all Google has done here is to make inconclusive accusations and use the occasion to highlight its own dominance over search.
It seems to me this is just a cheap and slightly seedy PR stunt.
Why? This seems like a great idea.
http://www.ft.com/cms/s/2/8b1ecaa2-cdb2-11df-9c82-00144feab4...
Would it have been better if Google had jumped straight to the questionable lawsuit part, like every other company seems to do when threatened on its own turf?
Those bogus results made it to Bing's results eventually.
Ok... So it proves Microsoft analyzes the toolbar behavior and when it has no other data, it will therefore look like a copy of Google search.
Sounds fair to me. Do you want to get into a discussion on how exactly Google tracks you online?
... in 7-9% of cases.
This is disgraceful attention-whoring on Google's part. Quite surprising, too, as I don't remember them ever stooping that low.
I wouldn't be surprised if they're more interested in data from domain-specific sites like epicurious than generic search sites like Google.
I'd also guess that this won't work once the SEO guys figure out they can feed fake clickstreams to MS.
They played dirty with Netscape/IE in the 90s and look what happened.
It is kind of embarrassing for Google, but if it is real and it continues, it's better to address it now rather than after MS becomes a titan of search and Google's market has eroded. At times, it seems like MS has changed in ways, but fundamentally they're still run by the same guys. Remember that when you play your Xbox or use Bing or any MS products, they don't like to see other successful software companies.
I remember that when Bing went out, everyone was wondering how close to Google the results were (and talked about it as a good thing).
In other words, there's a relationship between Page A and Page B if there exists a link beween them (==PageRank). But the strength of the relationship is increased based on how many users click on that link. I think that's the information Bing were trying to capture (or if they weren't, they should have been).
What I'm saying is that it's probably an unintentional side-effect. At scale though, the effect is that Bing gradually uses Google as a signal, simply because Google is a popular site.
edit: Yet another way of saying it: I think it's not just clicks on Google searches that are captured by Bing, but clicks anywhere. Google is a large site, so its influence on Bing can be measured. This is what we're seeing. My theory. I don't work in search.
edit: I'm not even sure if it's only search engines that are being analysed by Bing or all pages, but it's possible that it is just SEs - they could be capturing query terms distinctly.
Now the interesting thing to reverse engineer is what other information might be passed along to give relevance to the search term/click pair. If Google could establish that there was a third piece of info in the tuple, such as "originating search domain" and that Bing used this to weight term/click pairs based on the authority of the source, Google's claims would hold more water. I suspect that Bing has to apply some kind of validation of the term/click pairs (for instance, only sending pairs that appear on the same results page from accredited engines), otherwise they would be subject to "Bing bomb" attacks where users or botnets vote up lower ranked (or even unranked) clicks for a given term. (And if they don't validate or detect gaming, then there would be ample opportunity to inject all kinds of synthetic behavior into Bing's search results. Based on the relatively few number of users and clicks it took to own a long tail term, it seems like the protection they have is very weak or simple.)
I'm interested in another experiment. If you set up a honeypot, search for the term, but never click on the link, does the honeypot start showing up in Bing? The article doesn't say whether you tried this. Did you try it? Are Bing scraping your results from the page or only tracking their users clicks?
Google's search results are blocked in robots.txt, so I don't believe Bing has been able to crawl our search results directly. All the evidence points to users' clicks on Google, which are then sent to Microsoft.
Microsoft has (so far) declined to admit whether our allegation is true. Getting them to talk about exactly what they do and what software they use or don't use would be the easiest way. I'd like them to confirm or deny, which is why I wanted to go to this search panel later today and ask them.
Funny thing: http://www.google.com/search?q=site%3Abing.com%2Fsearch%2F
I was gonna call out Matt for crawling bing's search results but I'm guessing Microsoft hasn't realized they return results from the /Search/ folder. ;)
Isn't compliance with robots.txt more of a voluntary thing?
I'm not accusing MS of ignoring it when convenient, but if you/we/someone is accusing them of acting unethically wrt search results in the first place, telling the crawler to ignore robots.txt wouldn't be that far away, would it? (And likewise faking the user-agent, etc.)
For better or for worse, UA identification, robots.txt compliance - all those things are voluntary. I'm not suggesting they shouldn't be, but it certainly makes a difference in terms of whether something's possible or not. (And, if you ask me, places an even higher obligation on the actors to behave ethically, lest trust completely evaporates and the whole thing goes to hell in a handbasket).
It would take a pretty big leap to go from robots.txt is advisory to ignoring it constitutes a criminal action.
Again, if you actually read the article, you will come across the section titled "What About The Google Toolbar & Chrome?" I encourage you to read it.
[edit] Also, see this comment and patio11's subcomment further down the page, both of which were written an hour before yours: http://news.ycombinator.com/item?id=2165469#score_2165578.
I'm pretty positive that's not true. If you run Fiddler when browsing with Chrome you will see constant hits to toolbarqueries.clients.google.com whether you're using Google or not. I could be browsing some MS site and toolbarqueries.clients.google.com gets hit. Chromium doesn't do this.
Edit: You can uncheck everything under privacy and it will still send those requests.
Edit2: What it sends back looks something like this:
<?xml version="1.0" encoding="UTF-8"?><autofillquery clientversion="6.1.1715.1442/en (GGLL)"><form signature="8551191143090325242"><field signature="620769395"/><field signature="2995202485"/><field signature="2175865763"/><field signature="904516291"/><field signature="2953051246"/><field signature="2649047790"/><field signature="2308153337"/><field signature="1003471793"/><field signature="3255484099"/><field signature="1305698505"/><field signature="3676143819"/><field signature="1275502930"/></form></autofillquery>
Looks like auto-fill data, but this happens when I click around a site, NOT when searching Google or typing something in the address bar. For some sites (interestingly, not all) it sends 3 requests for each page load.
and this: http://code.google.com/p/chromium/issues/detail?id=60422
I would guess that Chrome is sending a hash of the <form> (perhaps URL + method?), plus a hash of each of the <input> tags, and Google returns some sort of information about what kind of form it is?
If so, it would mean it's pretty easy for Google to determine which sites you're on from the pattern of hashes sent for each site. e.g. I see this data sent in the clear for pretty much every page on https://www.facebook.com/
From the toolbar privacy policy: "Toolbar's enhanced features, such as PageRank and Sidewiki, operate by sending Google the addresses and other information about sites at the time you visit them."
Google has managed to demonstrate one way MS appears to be using the data. What does google do with their trove of data? That's a lot of data to collect and not do anything with.
If they want to make it perfectly clear they should add into their privacy policies and EULAs.
But the article clearly covers the available public statements on this issue and patio11 dug up a post from Matt Cutts in his comment below that directly addresses this: http://www.mattcutts.com/blog/toolbar-indexing-debunk-post/.
Please. Adding my own SSL cert to my own laptop is not harder than I'd expect. Certainly not harder than many other things you did in setting up this experiment.
[edit] I have always suspected that the real value of Bing for Microsoft is to prevent Google's data mining of queries originating in Redmond.
If I'm in the business of giving horse racing tips and I read your tips to see what your strike rate is compared to mine, that's one thing. If I start tipping the same horses as you, purely because you tipped them, that's quite another thing.
edit to expand: If the measure by which a search engine evaluated itself was advertising revenues, they'd all have massive intrusive adverts, and no users. The only viable measure can be the quality of the search results themselves. As a happy coincidence, if you build something capable of delivering high quality results, you can very easily use that to produce highly relevant adverts. Imagine that each advert is like a little webpage, and rank them just the same as you do for normal webpages. (caveat: there's no link graph for adverts, so we're reduced to using a simpler text mining approach, eg bag of words vector space la-di-da).
This whole episode points to the sort of counter-espionage operations the two companies are engaged in. Look how important a propaganda victory is for Google? It strains credulity to believe that the release of this information on the day of the panel discussion is pure coincidence.
MS are desperate to regain control. Google will soon launch their own web-centric OS properly, and bam, MS will have no business apart from selling to an ever-dwindling number of companies who can't believe MS don't rule the roost any more. In 20 years they will simply cease to exist if they can't come up with a world-beating online product and win back control of people's computing lives.
Notice how they're diversifying into games and search in order to prepare for the worst case; that their core OS and 'boxed software' business fails.
Why shouldn't Microsoft be using this kind of data? Google search result pages are part of the internet just like any other publicly available web site. Microsoft monitors what the users are clicking on Google and probably on Bing and other sites. So what? Monitoring users is not a new thing. It may be unethical and I may personally hate it, but almost everyone is doing it.
Google should stop whining about this and make their search result the best they can. If they had the best search engine, Bing could come close, but never overcome Google by just copying part of it.
It's ugly and immoral and probably legal. Good job catching Microsoft at it (and I think the really really unethical and scary bit is that Microsoft is cheerfully stealing info from users via their browser). I also realize that Google got where it is in Search by innovation and iteration, and that Google's search team has nothing to do with Android per se, but you might see how Apple people feel about the business empires built on stealing their ideas.
Then iPhone came out and it looked like this: http://km.support.apple.com/library/APPLE/APPLECARE_ALLGEOS/...
Now Android phones look like this: http://farm3.static.flickr.com/2795/4208849005_dd4b608729.jp...
You don't see any signs of copying here?
My question for you Matt... is there any way for Google to build a toolbar that effectively does what the Bing toolbar does (or even a joint one?). I jump to use the various search engines because no single search engine is sufficient. But clearly when Google isn't sufficient, you don't get the value of when I go to Bing. And vice-versa (as I don't use the Bing toolbar currently, but that may change now). Or do you feel that with 65% of the market, you don't need this info?
I bet is just recording general searches (input query) + clicked links. A pretty good idea.
And don't tell me you are not using the results from the google toolbar to rank the sites in google search.
The temptation to abuse that power is pretty big.