Why is Google so hysterically hypocritical about Bing using its public data?
roughlydrafted.com
roughlydrafted.com
Its easy to sum up. Yes, Google indexes hundreds of millions of sites. It does so in order for other people to be able to search and find those sites, which is important to their owners. Its a symbiotic relationship, not theft of intellectual property.
Google has spent billions of dollars in manpower and physical capacity in order to be able to do that. Also, each and every one of those sites can very easily say "don't index me" with a simple robot.txt file on their site. Some do, most don't, because most sites find it valuable for Google to index them.
Meanwhile, Microsoft is trying to compete in the same space as Google. They are also spending lots of money and manpower to build their search engine. But, by using Google's search results to improve their own product, they are acting as a parasite. Google gets no benefit from Microsoft using their data.
So to sum up my own summary: Google is a mass symbiote, Microsoft is a parasite.
What happened is that Google engineers engineered a use case, in which signal from those clickstream appear like a stealing of search results. Or in other words, Microsoft uses Google users behaviour in their ranking algorithm. But that's not unethical, and Google is doing the same thing.
If that leads to Bing absorbing Google's results, and eg. suggesting spelling corrections they would have never figured out except that Google thought of them first, then they are indeed stealing results, whether they meant to or not, and need to stop.
George Harrison didn't mean to copy the song "He's So Fine" when he wrote his song "My Sweet Lord," but he lost the lawsuit anyway and had to pay damages. Whether you call it "clickstream" or "user behavior," Bing is incorporating an association that could only have come from Google, and Google's robots.txt makes it very clear that robots are not allowed to mine search results.
> But that's not unethical, and Google is doing the same thing.
As has been mentioned many times, Google does not take user behavior from the toolbar as a signal for ranking.
Even if that's indirect information? For example, the Bing toolbar doesn't really say (in Gargamel's American voice for effect), "Nyhaahaha! I see these google search results! I will steal!" No, rather it says, "Ahh, after the result of a query to this site, we then leave the site to go to this page. If that query and this destination page appear in aggregate, Bing should take notice."
When I first learned how the Bing toolbar did this, I had two reactions: "Oh, that's clever. It's like engaging your users to help you be a directed crawler. It's 'querytext', which is not unlike google's innovation, 'linktext'." and then "But I would still never install the Bing toolbar. Man, that thing is a sad clown show." Honestly, it's one of Bing's smarter ideas.
There is only so far pure crawling and statistical analysis can go; there simply isn't enough data there and everyone knows how to use that data to great effect. Every search engine is incorporating new realtime communications and user behavior streams into their search results. Google certainly does this, albeit without a toolbar. Bing simply has the entire Microsoft software stack to lobby for help, so if someone opts into the Bing toolbar, they can opt into submitting additional information to improving the Bing index.
I'm not a big fan of MS or Bing, but in this they are only culpable for being clever. They are using querytext to improve relevance. I'm told you could make the same stunt work from a Wikipedia search box, leading to a wikipedia page.
I doubt they are only targeting Google. But I completely agree, the data is not being mined from Google, it's a relationship for search term a to page b created by the user, not stolen from Google.
As this comment points out, there is a decent argument that searchers on Google or other engines aren't the ones creating the associations, and therefore perhaps Bing has no right to use that data.
The relationship between a term and a page was not created by Google, it was created by users. Google just indexes everything and makes note of these associations, but its does not create the link in the first place.
Going down this road of debate, we'll be getting into the semantics of "inventing" vs. "discovering."
Turning data into ranking is the whole purpose of a search engine, just as turning data into theory is the purpose of science, and turning experience into art is the purpose of art. Essentially here, we're debating whether search ranking is more like science, in which there's a correct answer that you are uncovering, or like art, in which all of the product is subjective.
Having worked in the field for five years, I'd argue that it is far closer to art than science. Google's rankings are its subjective determination, and the courts have agreed that Google's rankings are its constitutionally protected speech.
Though the data may all ultimately come from the human-created internet, the transformation of that data is still important, and subjective. To claim otherwise is to miss the whole point of search technology in the first place.
I should think the answer to the above is 'no' in both cases, which is why this is cheating. It's probably not illegal, but I find the practice to be unethical.
Ask yourself this: If Google shut down today, would Microsoft be providing more relevant, equally relevant or less relevant results for those searches? If the answer is less relevant, then I think it's clear there's been a lapse of ethics.
In any case, for all you know, if everyone started using the Bing Toolbar, it may provide better clickstream data, causing more relevant Bing results.
Bing needs to blacklist Google from its clickstream. Simple.
That's one effect. While it's vivid, it might be a tiny side-effect only notable in contrived cases.
Overall, this kind of URL-after-URL signal, extracted from every participating user, and every trail through both search sites and non-search sites, might be discovering valuable terms-in-preceding-URL-to-later-visited-URL associations. These associations might result in many search improvements, other than the one-for-one result porting Google's experiment has found. We don't really know the relative magnitude of porting-results versus other-benefits, yet I think that's important to the analysis.
If a useful automated or user-driven process generates a little indirect infringement around the edges, is that enough to demand the process be stopped entirely? Note, that's not the standard Google wants applied to user uploads to YouTube, or excerpting of news and websites onto Google services. Google says: "defend yourself, by adding opt-outs (robots.txt) or sending takedown notices, and we'll undo the incidental infringements eventually".
The google engineers intentionally sent this click data to Bing, so is Bing really stealing? It's odd to act surprised when Bing uses the data that was intentionally sent to it. Bing could specifically ignore Google search results pages when it is tracking clicks, but is that legitimate? Google scrapes everything, why shouldn't Bing?
In fact, the data that it takes from Google engineers for carefully engineered corner-case searches is the exception.
edit: I should have clarified. I know that the Bing crawler likely respects robots.txt, but if they are using clickstream info to build their index, it seems right that they should respect robots.txt there as well, no?
But agreed that it seems like a very small nit to be magnified the way it has. Indeed, why they don't do this, and why we should care that Bing does, doesn't seem to be directly addressed other than by hand-wavery and PR speak about the research they've put into their algorithms and such.
* Are Microsoft saying end users are naturally clicking on nonsense phrases invented by Google?
* If not, what are they saying?
From googles mouth
http://googleblog.blogspot.com/2011/02/microsofts-bing-uses-...
"We gave 20 of our engineers laptops with a fresh install of Microsoft Windows running Internet Explorer 8 with Bing Toolbar installed. As part of the install process, we opted in to the “Suggested Sites” feature of IE8, and we accepted the default options for the Bing Toolbar.
We asked these engineers to enter the synthetic queries into the search box on the Google home page, and click on the results"
In all honesty I'm a bit wrong as they did say they used the google box on the page.
I guess I just have trouble seeing it taking more than about 3 minutes before 20 of the top industry engineers figured out a way to automate the process. Which is pretty long compared to the 2 seconds it would take for them to start thinking of ways to improve the SEO.
It's the human factor in Google's "experiment" that just doesn't fit. If they wanted a controlled approach, they would have written an application and run it and logged everything. Instead, it appears that they provided laptops so that the engineers could experiment and innovate their way to exploiting the Bing toolbar.
Thanks - that's what I've been getting at. This isn't data entry into the Bing toolbar, it's into a non-Bing page when one has the toolbar installed.
This is not correct. From Google:
"We asked these engineers to enter the queries into the search box on the Google home page"
Google did not enter the data into the IE search box.
Edit: I see you replied to yourself acknowledging the mistake - please ignore this then!
Got a citation for that?
http://googleblog.blogspot.com/2011/02/microsofts-bing-uses-...
If the Bing toolbar is picking up on this sort of thing generically (i.e. picking keywords out of the query-string on any page and associating them with clicked links, though I'm not sure how it could with a useful degree of accuracy in a way that couldn't be "maliciously" gamed buy underhand SEO activities) then I see nothing wrong in it as long as the users have knowingly opted in to their activity being analysed in this way. It would just be indexing keywords and content just as a web spider would.
If it is specifically detecting that it is on a Google page, and/or other search competitors, than the issue is much more cloudy.
Does Google agree their engineers did this?
This all happened in December. When the experiment was ready, about 20 Google engineers were told to run the test queries from laptops at home, using Internet Explorer, with Suggested Sites and the Bing Toolbar both enabled. They were also told to click on the top results. They started on December 17. By December 31, some of the results started appearing on Bing.
Sorry to be an ass. But you're wasting my time on this site, since I have to wade through your questions to get to the interesting ones.
* Google having the Bing toolbar installed and entering the search into Bing toolbar is one thing (and I'd expect MS to be using the data)
* Google having the Bing tolbar installed and entering the search into the Google hom page is a very different thing (and I'd expect MS to be using the data)
Judging from the moderation in this thread, people seem to think the first happened.
According to Google, it did not. No other source contravenes this.
Sorry if you think me pointing this out is bad. Perhaps your efforts would be better reporting all the non-hacker stories on the front page?
What happened to the Google line of "the more happy searching people we have online the better" (usually brought out in justification of their android investment)?
Shouldn't the same logic apply? Especially since at least half of the people using Bing aren't doing so voluntarily § shouldn't google be willing to trade that one search for having everyone happy on the internet and consuming more of it?
Is an incremental increase in Bing's search quality really going to take market share away from google? I'm skeptical.
§ 50% of Bing's traffic comes from fb, msn and windows live mail: http://techie-buzz.com/tech-news/google-is-4th-largest-traff...
How can Google opt-out of having its search index copied by Microsoft?
The funniest argument in this has been "Bing hasn't got the data, Google has an unfair advantage etc" as if Microsoft didn't exist before Google.
Microsoft have had years headstart to get their online division doing something worthwhile, but they still can't get it right.
What Microsoft apparently has done is to infer relationships from Gooogle SRPs to various sites by way of information gathered through the Bing toolbar (and possibly other documented and user opt-in channels w/in IE). Using toolbar collected signals is well-established practice -- that's how Google obtains information regarding site performance.
This is the crux for me, Bing is using data from Google, does Google use data from Bing or only themselves?
Even if all true, I don't see it as unethical business practice. Definitely a marketing black eye, but not much else.
Google is careful not to do this.
5.3 You agree not to access (or attempt to access) any of the Services by any means other than through the interface that is provided by Google, unless you have been specifically allowed to do so in a separate agreement with Google. You specifically agree not to access (or attempt to access) any of the Services through any automated means (including use of scripts or web crawlers) and shall ensure that you comply with the instructions set out in any robots.txt file present on the Services.
It's most certainly morally wrong.
I want to see a project started to r/e the bing toolbar, and create something to send rubbish data back to bing to screw up their search index. That'd be awesome to see.
They're capturing what users click on, which doesn't include any information about the wishes of the website owner.
I think this counts as extracting information from the Google results page.
<meta name=robots> is designed to prevent search engines automatically extracting information from web pages. By involving a human in the process, Bing are (presumably) avoiding this rule. Is this a good thing?
I accept that Bing don't access the page content itself, so wouldn't be able to see the <meta> tag if it was there.
It's dilger, why would you bother reading that tripe at all?
But he's wrong.
I can't understand why people make posts like this.
It's silly to hold onto some murky moral click-ownership here. It would still be silly, even if it hadn't given Google billions of dollars in profit, every quarter.
So if I have a site that review sites about x, google will use my links to rank the sites about x. It will benefit google, sites about x, advertisers for x, users searching for x. But NEVER me.
now, replace "me" for "google" and "google"for "bing". And you will see how you and google are being hypocrite
Here's a hypothesis. I suspect Bing is making use of data that is typed in the toolbar or browser's search box ... not data captured from text typed directly into Google's search page. This might explain why their "honeypot-take" rate was only 8%. This is a subtle point. So ... imagine if you are the coder who wrote this signal collection feature. Would you capture "term in search box,next URL clicked" OR would you capture "if search box search engine == BING or Yahoo or something else then capture term in search box, next URL clicked".
Let me propose an alternative experiment. The test clickers should have clicked on the second or third (or some position N where N > 1) link. This would have demonstrated if Bing is using the actual click information or the search results themselves directly. The former seems completely fair on Bing's part. The latter might be debatable. My point is this ... the Google geniuses fail to distinct these two cases and are muddling it up for the PR drones. This waste everyone's time and productivity. Moreover, they belittle the hard-work of engineers and scientists. It is sad that this is what it has come to.
From the official Google blog: We asked these engineers to enter the synthetic queries into the search box on the Google home page, and click on the results, i.e., the results we inserted.
Let's see:
Market cap: Microsoft - 239.49B, Google - 195.50B. Revenue: Microsoft - 66.69B, Google - 29.32B. Profit margin: Microsoft - 30.84%, Google - 29.01%.
Falling apart. Right...
"Daniel Eran Dilger is the author of “Snow Leopard Server (Developer Reference),”
quote down at the end. I'm sure his line of thought was: "All my friends at Starbucks have MacBook Pros, so Microsoft must be doing really badly."
Probably the same reason half his argument seemed to involve Android jealousy...
http://www.wolframalpha.com/input/?i=goog%20msft
Note how MSFT is swinging around 0% growth, while GOOG is steadily between +20% and +60% growth, in regard to share prices.
The market has spoken some time ago, and the trend doesn't look like changing anytime near soon.
Since a random point in time picked by WA it shows what you would be up on another random point in time. If you look at the historical returns of MSFT vs GOOG there is no competition. You're original $10,000 in MSFT would be worth many millions and your GOOG stock would be worth less than $100,000.
BTW., switch to the 5 year range in WA and you still see the trend.
They are looking in the worst shape to face the future than they have ever, in the last 20 years.
1. Future is also in the games, and here Microsoft has a very competitive product (xbox/kinect)
2. mobile race is not over yet, and WP7 is conceptually a good product, combining polished UI (aka iOS) with at least some hardware variety (aka Android). It just came late to the game, but it is not yet clear if too late. I personally think they can catch up with throwing enough money into it.
3. We will see when they release OS for tablets. They are late here as well, but similar to point 2), maybe not too late.
I would say that WP7 has a superior conception to iOS, and I am a hardcore iPhone user. My phone is really about my communications and managing my data away from home. Apps are only a means to that end. Right now iOS is pretty App centric. WP7 on the other hand is personal data centric, which in the long run will provide better functionality for "managing my data away from home" with less friction.
I wouldn't want to be in MS position, but with enough money, patience and perseverance you never know what they might do and what may be next
Apple likes to retreat into it's niche where it controls everything (helps avoid antitrust regulators breathing down its neck) and has high margins. This is already happening with the iPhone. Android while gaining in popularity, lacks the polish in the UI and performance.
Even with much better raw hardware, Android performs worse compared to Windows Phone 7. See http://pockethacks.com/windows-phone-7-smoother-than-dual-co...
The going empirical formula is version 3 of their products are when they start becoming good. While WP7 is technically version 7, it's actually a complete reboot of the platform with zero backward compatibility. They're taking the middle route between iPhone and Android, having variety in hardware but strict hardware requirements to prevent fragmentation and a controlled app ecosystem. That said, they don't have a good tablet strategy. Maybe Windows 8 will really be tablet optimized, or maybe tablet hardware will catch up to make it run smoothly(the Win UI is still a hard problem on tablets though). Or maybe they will port the Windows Phone UI to tablets.
In short, very interesting times are ahead in the mobile/tablet space and it would be a mistake to write Microsoft off.
But it's pretty clear they lack the leadership at this point to move quickly to remedy their shortcomings. That wasn't true at the come-from-behind points in the past (Gates was still around).
<quote>Shame on your pretentious, obnoxious, indefensibly egregious double standard in the field of using public information to turn a profit.</quote> <<-- Ironically, Bing is doing the same and unlike content creators who can opt out by putting a robots.txt file Google cannot opt out of this because like it or not toolbar is always sending back information.
Google indexes public content but how it aggregates it and displays it is their deed. The ordering, etc is Google's work not content creators.Copying it is just like copying an IP. I would love to see this go to court and be fought.
Google makes money aggregating other people's content. What happens when people aggregate Google's content? What's fair?
Google are objecting to Bing seemingly passing off Google results as Bing's own (and not including their ads, no doubt).
It's clearly legitimate for Bing to index a Google results page if it follows a link to it (assuming absence of robots.txt etc) and visa versa.
Does it matter that Bing obtained the Google results via a search from a human not affiliated with Bing?
Does it matter that Google blocks crawlers from accessing its search results, using robots.txt?
Does it matter that Google's result pages don't include the noindex instruction?
Tbh, if I was Google I'd mark my results as noindex and see whether Bing respected that.
You say Google adds value. Other's don't think so.
Why can't I add value to Google's content and present it in my own way?
Well now people are adding value to Google's products and presenting them in a new way. This opens up a can of worms for Google.
Google collects data too, Google does its own shit like everybody else, like Microsoft. Google has its own toolbar too. And all this story seems more a marketing move against Bing.
What I want now to happen is to have more competition, instead of crying Google should work hard to be unbeatable and competitive. History teaches us a copy it's never better than the original.
It's Spy vs Spy:
http://en.wikipedia.org/wiki/Spy_vs._Spy
...a pitched battle between the opposite-but-indistinguishable agents of two superpowers, on a plane so far removed from the realm of actual people and their concerns that the customers don't even appear in the frame.
Arguing that according to the strict letter of the spec, robots.txt only has to be obeyed by true crawlers does hold water as strictly speaking Bing could disregard robots.txt altogether - it's not certainly not legally enforceable. The intent of robots.txt is clear and Bing should be trying to obey it wherever possible.
In this case it appears the keyword was associated because it was typed into the Google search box, and from the there to the faked destination. IMHO when the clickstream was analysed it should have disregarded clickstreams that pass a robots.txt-excluded page as these could establish associations that were not supposed to revealed by crawlers.
We aught to hold Bing and any search engine to the highest standard when talking about crawling etiquette.
If I'm a site that blocks /dynamic because I want to make sure my stats are accurate, not factoring bots, and I want to keep the load down on slow-generating pages, then I'm perfectly happy with the data being collected.
Is it the intent of robots.txt itself to block the clickstream data, or is it just google's intent?
edit: bing toolbar is effectively an extension for MSIE. robots.txt is designed to stop spiders. Spiders aren't extensions. Therefore that's not what robots.txt is for. What did I say wrong to get downvoted?
I think that's what most people though was wrong in your comment.
Care to elaborate on how an adblock extension could be excluded via robots.txt in that fictive scenario?
User-Agent:adblock
Disallow: /
Or even, if the "User-Agent: (star)" is supposed to tell Bing Toolbar not to play about with the page, then why doesn't that also apply to all your other extensions/addon/BHO's/plugins whatever you want to call them?If Bing uses Google's search results, then we effectively only have one major search engine. Yahoo has switched to the Bing engine, and Bing gets a big chunk of it's results from Google. So we have a search engine monopoly. That's wrong with that.
> "This is the company that indexes blogs [...], and then makes all this information available without consent"
He first asks what's wrong with Bing copying Google's results, because it's public content, then says it's immoral to index public content. Double morals? Logic fail?
Google complains about Bing using public content. Google uses public content.
Snarky question? Fail?
Nerd fight.
Not a mention of GoogleBot.
Yelp's API specifically says their data cannot be aggregated with other sources. When presented with that, Google simply said they weren't using the API, they were scraping.
Apparently scraping is a free pass to do whatever the hell you want if you're Google.
However, please avoid linking to roughlydrafted.com. At the risk of going ad-hominem, I just wanted to point out that if you have read any of the material on the site, you will know that it is just full of flamebait articles with an absurd Apple fixation.
Everything Android does is copying Apple. Everything Microsoft does is wrong and stupid. If flash is not supported on iPhone it is because it is the morally correct thing to do. When Apple announced the iPhone without an SDK roughlydrafted called developers stupid for asking for an SDK and that javascript on the web is the new SDK. This post however takes the cake, you must side with Bing because Google is a greater threat to Apple than Microsoft. Seriously, how does Overture even enter into this discussion. It is tiresome to see this site linked to all over the web.
robots.txt and noindex http://www.google.com/support/webmasters/bin/answer.py?hl=en...
And "Install the Google Toolbar and do a search of Bing, and Google actually directs your clickstream back for its own analysis"
Google/Matt Cutts have been very open about what they do and don't use that information for, and they aren't using click tracking.
And there's truth to it.