Bing: Why Google’s Wrong In Its Accusations
searchengineland.com
searchengineland.com
Imagine this scenario, I create 10,000 VM instances of windows running IE8 with the Bing toolbar. I create a local host to 'stand in' for Google such that it emulates the actual Google site (one could even used scraped Google content) and for my 'target' query it returns my spammy results, and then my VM machine clicks on one of the spammy links.
It seems that Google's sting worked because their queries had small return rates, but with some resources it would seem a viable way to inject SEO love right into the bloodstream of Bing's ranking algorithm. As I see it, a whole new front just opened up for spammers.
--Chuck
The battle against spammers will continue to be cat and mouse game. And when one of your signals gets exposed, of course it will be a target.
There's nothing new here. In fact, I think most spammers already assumed that the toolbar clickthroughs fed into both Bing and Google toolbars. I did.
If the spammer gets 10,000 IPs then more power to him. But there's probably more cost effective ways of black-hat SEO.
In short, this gives them a new service to sell when they rent out their existing botnet.
Or, what's more likely, is that the respective companies are aware of this very obvious vector and have attended to it.
Botnets are run off of the computers of ordinary, clueless folks, who might be real Bing users submitting real data in addition to whatever the botnet sends.
I've already linked to an analysis of the actual protocol data submitted by the toolbar and I can see obvious ways to copy and fake it.
If you use only the computers in your set that already have Bing, copy that unique ID and figure out what IP they're sending it to, your data will be identical to that sent by a real user.
At that point, you have to harden the protocol and hope it stands up to reverse engineering, or start spam filtering it (if they weren't already). Maybe they can do a good job of that, but it really lowers the quality of the data they're getting once enough people are feeding them garbage.
SEO types already set up thousands of spam websites to game PageRank. I don't see why this would be any different.
Of course my favorite was the Amazon hack of repeatedly putting some item in a shopping cart and then adding in a bit of pr0n (or some other weird combination) so that Amazon would put the other item in their auto suggest product combination.
http://projectgus.com/2011/02/bing-google-finding-some-facts...
I'll save you some reading and point out that this is the important part, the exchange with Bing:
http://projectgus.com/files/googlebing/seaport-trace.txt
I bet I could duplicate that with a few lines of Perl. Then I could feed Bing whatever I want, no VMs or clicking or DNS/hosts files to worry about. You might have to harvest a valid identifier (such as the one linked there) or do some figuring out so that you submit this to the right Microsoft IP.
If I were working for Bing, I'd start looking at hardening this a little, before the SEO people figure it out.
"PR is not leading this dispute. It’s following behind. This dispute is happening because real engineers at Google felt there was a deep injustice going on — as reflected in the quote from Google’s Amit Singhal in my original article. I’ve known Singhal for years. I’ve never seen him speak like this before. It’s not because Google PR told him to. It’s because he’s fundamentally bothered by what he’s seen — as are members of his team.
This dispute is also happening because real engineers at Bing feel there’s a deep injustice going on — as reflected in the quote from Harry Shum above. Bing’s worked incredibly hard to build a search engine that’s worthy of respect. Now here’s Google suggesting that Bing has simply cheated its way to relevancy."
The conversation with moultano on a thread a couple days ago was a good example of this. http://news.ycombinator.com/item?id=2177354
"We’re not going to stop using that signal, unless it messes up relevancy. It doesn’t make sense to exclude that large amount of traffic from our usage set," Weitz said.
From the TechCrunch article a few days back [1]:
"Google had employees log onto ms customer feedback system and send results to Microsoft."
(to which Matt Cutts replied: normal people call that "IE8")
Unlike many others, I do not think this is a cut-and-dry issue, but the squirrelly responses from the MS folks on this have really made me think they are just up to no good, true to the form of so much of their corporate history. Arguing with vague technical terms and ad-hominem attacks are not a good way to convince a highly technical crowd of your virtues.
And Google's sudden silence you don't find suspicious. This is the same company that invited the author of the original article to their headquarters the day after he wrote it.
2) More importantly, why would Microsoft need to prove google is also using clickstreams? Microsoft does not believe using clickstreams is wrong.
If they can make Google a hypocrite, the issue goes away. that's motivation enough for Bing to investigate the Google side (though I doubt they'd find much).
http://www.google.com/analytics/tos.html
The 'privacy' link from Analytics leads back to the general privacy pages, and their overall privacy policy says Google may use any information they have (from logs, cookies, etc.) to "[p]rovide, maintain, protect, and improve our services (including advertising services) and develop new services".
I can of course say, I use this feature, and of course point to others who do, but I don't have data to say it is significant. However, you have stipulated it is not, but have not provided any data to prove it.
there's no way they are just magically doing this for "every" search box on the web.
"Here’s another one. This time, it’s a misspelling of “bombilate,” a rare word I cited above. I searched for “bombilete,” instead.."
In essence, they say that Google only pointed to the typo, but Bing redirected to it. Thus, "it’s very unlikely it figured this out from Google". For me, making that argument is insulting to the readers intelligence.
"But it’s very unlikely it figured this out from Google, given that for the misspelling, Google doesn’t auto-correct the word nor provide the same answer."
Bing's nr 1 result is not in Google's results at all.
It sounds like they automated the process of piggy backing off the work of other search engines, not just Google.
They should exclude all competing search engines from this process.
> Well, above is the same situation where Bing gets a misspelled word right — a link to a definition of the correctly spelled word at the top of the list. But it’s very unlikely it figured this out from Google, given that for the misspelling, Google doesn’t auto-correct the word nor provide the same answer.
Perhaps because people first click on the spell correction and then on the result - so maybe they don't yet copy also the spell correction, only the results. I think this is even stronger that they copy the results, not weaker.
Google said in October that it found statistical evidence that Bing suddenly became more Google-like. More listings in the first page of results of both search engines seemed to match, as did more of the number one results.
How would you get statistically significant results for such things, over time, without constant automated probe queries against Bing?
I think such probes are both legal and wise... but Google should drop the pretense that robots.txt is a sacred barrier across which no analysis can be done, no matter how indirect or for what purpose.
Also, I'd wager at some time in its history – if not constantly even today – Google has shown panels of users results from Google and its competitors in various combinations – side-by-side, with and without branding, intermixed randomly – and used their reactions to detect areas where the competitors are doing well, and Google could improve.
Further, either human eyes or algorithms then tried to determine adjustments to close any gaps in user satisfaction. The net effect of any such process is – surprise, surprise! – leveraging strengths of other engines to patch weaknesses in Google. This is normal, expected behavior by any serious search competitor.
It is very very different. You can't tinker with the weighting of something that doesn't exist, and if you don't have the necessary data, no amount of reweighting is going to improve things. Microsoft in this case no longer needs to come up with any data of their own, they can just use clicks on Google as a proxy for any combination of signals.
Amit has made it very clear however, that we will never do anything that would cause us to directly or indirectly copy a competitor's search results.
But given what Googlers can't say, and don't even know about what other groups within Google are doing, it's not 'very clear' to me that you aren't already doing very very analogous things, with regard to every other site on the internet.
You've got the data; you're allowed to use it by your privacy policy; you've got the rationalizations handy. ("Sites didn't block us; fully-informed users opted-in; this is a crucial way to fight the manipulators; it's only helping us weight things we already found by other means; etc.")
Amit's not made anything clear to me, with his finessed "put any results" wording. Danny Sullivan picked this up too, as he remarks in the headlined article:
Google’s initial denial that it has never used toolbar data “to put any results on Google’s results pages” immediately took a blow given that site speed measurements done by the toolbar DO play a role in this. So what else might the toolbar do?
There's wiggle room in the definitions of 'copy' and 'competitor' in your 'never' promise, too. Is it OK if Google Toolbar data hoovers up implied editorial-quality signals from user navigation on every site that isn't a 'competitor'? (And given Google's size, what site isn't a competitor in some respects for audience share?) Is it 'copying' if your use of clicktrails makes a preexisting result move from #11 to #9 after you observe it satisfying people in other browsing sessions? Move from #99 to #2?
(Has the effect of any of Google's competitive analysis ever resulted in a single result moving closer to the position, higher or lower, that it had at a studied competitor? Some people could call that 'copying'.)
Maybe none of the clickstream sources Google uses stick out as a dominating factor because Google has so effectively "commoditized its complements" – and no one other entity (except maybe Facebook) has access to as much clickstream data as Google does, simply from its own sites.
Given that, it seems a little convenient that Google's standard is "every aggressively creative use of behavioral trails that led up to our 70%-90% share dominance was OK, but from now on let's be really rigorous about letting others observe our info-effluents."
Speaking only for myself and only on the ethics, I generally feel that any site that allows itself to be indexed is pretty happy with Google (or bing) doing whatever they can to rank it better. Even with the link data that sites provide, you can add rel=nofollow to the links if you don't want search engines to use them but still want your pages indexed (yelp has done this for instance.)
For me that's the ethical boundary. Sites have various ways of indicating their wishes, and that ought to be respected in spirit beyond the technical details.
Legally, the technologies that make the internet work all rely on the idea of fair use, so it is very important whether something is "fair."
If Google doesn't want IE features or the Bing Toolbar observing its site interactions, it can disallow such visitors. A steep price to pay, at too coarse a level of control? Yes, just like a site deciding to bar Googlebot.
I would agree that a 'fair use'-like analysis makes sense.
I would further agree that any site solely, or predominantly, powered by indirect observations of Google users would be an unfair taking. You'd crush such a site in court.
Meanwhile, a site that tallies Google referrer inclicks for itself, or for a network of participating sites (as with analytics inserts), even republishing summaries of Google source URLs and search terms as public data, is almost certainly fair use. It's taking data you're dropping freely onto third-party site logs, and making a transformative report of it.
What Bing is doing seems to me somewhere in-between. The mechanism avoids literal copying of specific artifacts but the net effect in some cases approaches the same result. As with other 'fair use' analysis, it's rarely black-and-white. The magnitude of the information used, its effects on the market, and the value-added transformation afterward are all important. I don't know how a court would rule in such a suit but the discovery process would surely be fun for spectators like myself!
Shrum:
Not so, Bing told me. In October, Bing says it rolled out a new
ranking algorithm plus a new experimental system called “Aether”
that allows them to test changes in their ranking methodology.
That’s what caused the bump that Google saw, not some sudden use
of the surfstream, Bing said.
Web (during October 2010):http://www.google.com/search?q=Bing+new+ranking+algorithm...
For example, if you did a search on Amazon, Bing might detect that. A search on eBay might get spotted. A search on Yahoo, that also might get extracted. Any number of searches might be identified. Bing would associate the next page you went to after doing those searches as being a possible 'answer' to those searches.
It's a much different case when Microsoft is copying search results from Google or another whole-web search site. There Microsoft is directly competing with their product, and the bottom-line is that it's fundamentally fishy for Microsoft to be using their own search results against them.
Somewhere in the world at this moment there's a person with the Bing Toolbar installed searching for something using Google. One of his results is to a website, for the sake of narrative lets suppose it's one of our own: A YC startup.
The searcher clicks to the YC startup, it's exactly what he's looking for, he converts.
Now, another searcher, using Bing by choice, searches that same term. That YC startup is in the results. The click is made, another conversion happens.
The first user consented to his click analysis by installing the toolbar.
The startup will surely agree that yes, we are a very good result for that term! We should show up on Bing, DDG, Google, Whatever! And we don't care how we get there!
The second user gets a result that's maybe ranked higher because of the first users click was noticed by Bing.
Google doesn't lose a customer because the second searcher was already on Bing to begin with.
Bing does nothing but analyze their own users' behavior (users of their toolbar) to deliver better results.
I find controversy over this beyond absurd.
The toolbar says nothing about taking his data to improve Bing, only for Site-Suggestion which, to average user, is clearly a browser feature.
> Google doesn't lose a customer because the second searcher was already on Bing to begin with.
What about one who could have converted to Google if he search on Bing and didn't find a result?
Google couldn't lose a customer but there's no way Google would gain a conversion from Bing in this situation.
> Bing does nothing but analyze their own users' behavior
Assuming that this is a search result that Bing wouldn't have found by itself, then this "user behavior" wouldn't have occurred for Bing to capture had Google not exist.
Without Google each user can then probably only have behavior of clicking same old web site they have collected from long ago, there's no search engine for them to learn new relevant site easily.
But the issue is, my clickstream is mine. If I chose to share it with somebody -- Bing, Google, My Mom -- it's my choice.
If I, as a user, don't want Bing to have access to my clickstream, I won't install their toolbar.
But if I do, it's not Google's business.
So, it looks like bing is indirectly using google's data by incurring lesser cost and this is how they seem do it.Spy on google searchers and build the right database. instead of spending on innovation and new ways on the algorithm. Let google do that part while we piggy back on their good ones.
This is what has really irritated google.
These battles do nothing more than establish the two dominate players.