Bing sets the record straight on recent accusations
bing.com
bing.com
"This [torsorophy] example opened our eyes, and over the next few months we noticed that URLs from Google search results would later appear in Bing with increasing frequency for all kinds of queries: popular queries, rare or unusual queries and misspelled queries. Even search results that we would consider mistakes of our algorithms started showing up on Bing."
That's the key. Not the honeypot. The honeypot was just a test to see if they could catch them red handed.
Bottom line is that Google looked at the statistics and found that Bing results were improbably similar to Google results. Maybe they're lying about this, but I doubt it.
There's not enough information to figure out exactly to what extent Google impacts Bing results, but I would bet a lot of money that in the Bing code base there is Google specific code that behaves along the lines of "if on google, do x" rather than some generic code that just targets all sites.
Yes, Bing is likely using Google's data (if available, via Bing toolbar) as part of the ingredients. You can argue whether this is stealing or not, whether this is Google results or user search/visit behavior, and but I would say it isn't exactly "if on google, do x".
I wish bing had taken this opportunity to answer that question in this blog post. Since they didn't, it makes me more likely to assume the worst. All they would have to say is something like "We use anonymized click data and don't special case Google" to avoid most of the controversy. Then we would be back to debating whether it is ethical to do something theoretically ok even if you know that it practically means you'll be cribbing from a competitor.
Just call it "imitation is the most sincere form of flattery" and be done with it, would you, Google? Spend your time improving UI, defend your cash cow against spam, focus on making Android better, be more coherent in your social/location strategies, lead the industry in privacy (not just talk of privacy). All of these will benefit users more.
But the argument "I rather Google spend.." can be applied to anything - it isn't as if they haven't be tackling other problems, and they are clearly worried about bing (at least in PR, bing spends big - lots of MS employees posting on forums, including here, which I am sure is encouraged). Bing is losing a lot of money for Microsoft, so by that token, I could say that they should stop it, and spend money on xbox (which is fantastic - no idea if it is losing money, but what a leader) - make it stupidly cheap, make it dominate the home etc ...
Why should Google improve their search spam algorithm? Would they gain anything if MSFT then their compares their own results against Google's, to delete the spam farms thanks to Google's work?
How do they look at what a user searches and what result they get? Ok, the "click stream" can see what pages a user visits.
But to get the search term, they are parsing something, and unless they implemented a universal search recognizer that will rank up results from any old site's search (allowing 20 SEO guys to push private SERPs to the top), it seems more than probable the parser indeed would start with "if on google, do x".
referrer in the http request?
Hand waving that this is "clickstream" data doesn't mean Bing is not looking specifically for Googled search terms and the resultant user selected URLs.
I personally don't mind them doing it. I just think their post is non-responsive in hopes of persuading the non-technical reader there's "no deliberate copying to see here" when they're clearly intentionally ingesting an extraordinary volume of Google-keyword-to-Google-result data and using that data to map keywords seen only on Google to results on Bing.
Flattery, etc.: http://www.wired.com/epicenter/2009/06/kayak-bing/
* When query Q gets made more than N times at bing.com,
* Mine clickstream data for the next M urls requested after searches for Q
* Any url that appears more than T times (possibly spread across some number of users) is presumed to have been found relevant to Q, and derived either from later searches (corrected spelling) or curated sites or other search engines. Add to mapping of valid responses to Q.
It's not a very good algorithm, of course, and if you have any other source of information about Q you're probably better off using that instead. But it or something like it could explain the torsorophy example and every other part of Google's narrative, and it's not particularly suspicious or questionable, and it certainly doesn't involve targeting Google.
Put another way, it's impossible not to clue Bing in on at least the fact that you are making these searches.
From http://goo.gl/Bi0JH (Google blog):
"We gave 20 of our engineers laptops with a fresh install of Microsoft Windows running Internet Explorer 8 with Bing Toolbar installed. As part of the install process, we opted in to the “Suggested Sites” feature of IE8, and we accepted the default options for the Bing Toolbar.
We asked these engineers to enter the synthetic queries into the search box on the Google home page, and click on the results, i.e., the results we inserted. We were surprised that within a couple weeks of starting this experiment, our inserted results started appearing in Bing. Below is an example: a search for [hiybbprqag] on Bing returned a page about seating at a theater in Los Angeles. As far as we know, the only connection between the query and result is Google’s result page (shown above)."
1) When a user enters google.com/search? , scrap the page (pay special attention to their spell corrector)
2) Send all the data back to MSFT
3) Profit nicely!
It's simpler, and probably works as well as the best competitor out there ;)
Google need to go to some extra steps to show that Bing is copying Google links for popular terms. If Bing is weighting clickstream data from Google searches very highly, that is more or less admitting that other search engines work better.
If that is what's happening, then the moral/ethical argument against Microsoft would have to be that they should treat Google specially, by explicitly ignoring clicks on Google's result pages. That seems to me to open up quite a can of worms.
If anybody's search toolbar checks a site's robots.txt before sending clickstream data, I would be very surprised.
A client-side robots.txt rule would also make anti-phishing features trivial to bypass...just put a robots.txt on your phishing site.
A useful post would have addressed at least 3 things:
- What are the specifics of the mechanism by which they ended up obviously copying google's results?
- Do they handle clicks from google differently than any other clicks?
- How different would their results be if they didn't use clicks from google's search as a signal?
Additionally, if things were reversed and Google was posed these questions, I would imagine would be lacking just as many answers as what Bing is supplying.
There is a subtle point that many posts like this are overlooking. Google didn't run this experiment to cause bing to show bogus results, they did it to confirm the rise in suspiciously similar results produced by bing.
Google didn't run this experiment to cause bing to show bogus results, they did it to confirm...
There is a counter argument to that. If Bing's claims are true, then Google didn't run the experiment to confirm the results, they ran it to cause the results.I think 99% of HN, including myself, will share your opinion, but it doesn't counter Bing's claims because it boils down to "which company do we believe?"
Bing: doh, of course we use google results, but we told you that before in an obscure academic paper. Didn't you read that? And we are innovative - shiny pictures!
That's like saying you'd be evasive like OJ Simpson if you murdered a few people and were on trial.
So if Google used Bing to rank searches it would also be under scrutiny? Since that's not the case we can all avoid the useless thought experiment can't we?
Chrome tracks clicks and traffic. Google Toolbar tracks clicks and traffic. The difference in the situation is that the majority of people use IE to search on Google, not the Google toolbar/Chrome to search on Bing. Bing has more data to work from than Google does in regards to this exact type of click tracking.
That's the important one. These Bing people are happy to run their mouths about 1k signals blah blah blah. Who gives a fuck? What is the weight on those signals.
Further, this case was made obvious because Bing couldn't create an answer to a query so they copied Google's answer. Dude admits as much. Why should we give Bing a pass on copying Google's results just because in this instance it was too hard to find their own results?
Low enough that it took Google nearly three engineers per successfully injected honeypot (7 honeypots per 20 engineers) and Google was only able to achieve a 7% success rate despite their extensive in-house knowledge of SEO.
Where are you getting the 7% number?
That sounds reasonable.
Again, that's not the allegation, that's the evidence. The allegation is that bing is using data that essentially amounts to a wholesale copying of google's results.
Well, if there's one source of data, that data is weighted at 100%.
The whole experiment is meaningless.
- Do they handle clicks from google differently than any other clicks? Hopefully they do if they are using user clickstream data. Each domain should have something like a page rank to determine how trustworthy it is.
- How different would their results be if they didn't use clicks from google's search as a signal? Why would they not use Google. Google is a site on the internet where users click links, Microsoft is collecting data from sites on the internet where users click links.
If there's some algorithmic reason to believe that Google gives more trustworthy results, that's one thing, and the pagerank weight can be empirically determined. If, on the other hand, there's some point where they explicitly treat google clickthrough data any differently (if site=='google.com' weight+=10), then that seems to have crossed a line.
I'm trying and failing not to have this sound like some kind of koan, but if you treat everyone differently in the same way, then that's (arguably) fine. It's only if Google gets some specific attention that it seems more malicious.
Obviously they won't explain the first one, just as Google won't explain the specifics of why StackOverflow scrapers sometimes outrank StackOverflow, and for good reason. Explaining the specifics of the mechanism would open it wide to spammers.
The second? Yeah, that's the million-dollar question. They should answer that. Definitely. It touches on the smelliest thing about the issue. If I had a search engine (I don't), and I wanted to copy Google's results (I wouldn't), and I had the ability to collect user click data, I could use that click data to create plausible deniability for the copying. This is exactly the sort of thing Microsoft has done in the past on other issues.
The third? It's certainly relevant, but I doubt Google would be willing to tell me how its results would be different if it didn't use a specific metric.
> Google engaged in a “honeypot” attack to trick Bing. In simple terms, Google’s “experiment” was rigged to manipulate Bing search results through a type of attack also known as “click fraud.” That’s right, the same type of attack employed by spammers on the web to trick consumers and produce bogus search results. What does all this cloak and dagger click fraud prove? Nothing anyone in the industry doesn’t already know. As we have said before and again in this post, we use click stream optionally provided by consumers in an anonymous fashion as one of 1,000 signals to try and determine whether a site might make sense to be in our index.
Beyond that, it's kind of crazy to think they'd open up on nitty-gritty details about their algorithm - nobody does that.
Anyways, I'm still with Google. I hope Google wins. But attacking Bing here was probably a tactical mistake - pretty much all marketing thought ever says "Don't attack-market against upstarts if you're the market leader!" You can't win if you're #1 and you do that. Google's #1. Attacking Bing was a really bad tactical move, though I'm still casually rooting for Google to win.
hunh?
> So big and noticeable that we are told Google took notice and began to worry. Then a short time later, here come the honeypot attacks. Is the timing purely coincidence? Are industry discussions about search quality to be ignored? Is this simply a response to the fact that some people in the industry are beginning to ask whether Bing is as good or in some cases better than Google on core web relevance?
I can't believe this was written by someone with 'Vice President' in their title ... how childish.
Google fell into the classic trap of confirmation bias with bad scientific method. And thus the test was 'rigged', they never had a control variable.
Wouldn't it be more accurate to say that "bing is using click data"? The fact that google.com is in a lot of that click data is a questionable decision and the root cause of all this drama.
The only way the words "Bing copies google" would be justified is if MS were directly querying Google on certain keywords and ripping off search results. Google have provided no evidence to suggest this. I expected the commmenters of Techcrunch to be unable to grasp this, but it seems that HN is often like this too.
It would even be justified if they were harvesting click data only from Google (or explicitly treating that click data differently), because then they're just doing the last one, but obfuscating it: it would be like refusing to bribe a politician directly, but instead making a large "investment" in a corporation they own.
It's not a justified claim if they built a mechanism that genuinely gathers interesting data, and would continue to do so in Google's absence. I think Microsoft is claiming this, but their responses have been so murky that it's not 100% clear. Google certainly hasn't produced evidence that renders this version implausible.
If you really want examples, just watch Fox News or MSNBC for fifteen minutes. You'll probably see at least one or two examples in there somewhere.
"Bing results increasingly look like an incomplete, stale version of Google results—a cheap imitation"
Which they have NOT demonstrated. Their results can easily be interpreted that they imitate user clicks.
That is not the hypothesis. The hypothesis is "Bing has results in its index that it could not have gotten in any other way than from Google search results." Their experiment does indeed confirm that hypothesis.
The inverse of the conclusions of their experiment are also incorrectly assumed (Google results, using IE8/Bing toolbar, make Bing results != Bing results, when using the IE8/Bing toolbar, are from Google).
What is that scenario?
They never tested the fact that it could have come from any other website as well. Thus, they can't conclude that Bing is copying Google or whether its copying the user's browsing behavior.
That's all Google is saying.
> They never tested the fact that it could have come from any other website as well.
It doesn't matter, Google is only complaining about what Bing has copied from Google. What Bing copies from other sites is between them and the other site.
The fact is that they used unique search terms that link to unique subjects. They ONLY submitted searches using what they described. That means the only places those searches passed through were the OS, toolbar, IE and Google. They know Google received those searches, did it propagate to Bing?
That's it. They didn't need to offer a placebo to anyone. It's like throwing a ball and hearing an echo. You don't need to NOT throw a ball just to make sure it doesn't echo.
That's about as unclear as anyone can be. Seriously Microsoft, this is the Web. When you want to point your readers to a document, don't fucking cite it like some printed academic paper! LINK TO IT.
The yearly subscription is $1500.
If [2] can be transposed, non-subscribers can also get the report for a whopping $750.
I couldn't find that report in their archive,though. The only article published on june 15 2009 is [3], which is unrelated.
[1] https://www.directionsonmicrosoft.com/
[2] http://www.reuters.com/article/2009/04/21/idUS223966+21-Apr-... :
[3] http://www.directionsonmicrosoft.com/update/50-june-2009/678...
The Bing toolbar uses clickstream data to extract information about the relatedness of urls, just like any search engine crawler does (page rank works this way). When a user clicks on url B from url A the Bing toolbar sends information back to MS about the relatedness of urls A and B including all the meta-data involved. If url A is ".../search?q=torsorophy" then Bing will make a note of that and will start showing results for B when you search for "torsorophy".
In principle this isn't necessarily a bad thing, it allows Bing to index sites that it wouldn't otherwise, letting its search index grow more organically. However, when search engines come into play things get problematic, because now the Bing toolbar is little more than an automated method for scraping search results piecemeal. Given that a great deal of modern web surfing falls into the "search for X, click links for X" pattern, this should have been something that the Bing engineers anticipated (if they didn't anticipate it that's bad enough, if they did and ignored the problem that's much worse).
Worse yet, the Bing toolbar is effectively a search indexer which does not respect robots.txt. Let that sink in for a while. Google.com/robots.txt has this line: "Disallow: /search", and yet apparently the Bing toolbar has absolutely no compunction about effectively ignoring that.
tl;dr: MS has created a search indexer which ignores robots.txt, this is bad.
The user gave permission to Microsoft to use this info by installing the toolbar. I don't see why or how the Bing tool bar should visit Google.com/robots.txt to see the blocked folders unless they were crawling Google pages.
In particular, you copy the most popular results, so if you weight the % of google's index that you copied, then it's going to be a huge number, and almost for free.
It's evil. But devilishly smart.
Personally I don't think that flies. Should I be able to create a toolbar that rips the data from hulu.com and hosts it elsewhere?
maybe i don't understand, but your logic seems correct until the above statement.
suppose you had a system that only used clickstream data. so you store a big list of url pairs (A,B) and a probability that B will follow A. i believe that your argument relies on the fact that it's possible for this system to violate a robots.txt file and i don't yet see it.
This could easily be fixed, by checking the clickstream data against robots.txt files and discarding data that shouldn't be used. Microsoft apparently has decided not to take that step.
- the "intention" of the robots.txt standard is as you state
- the url is included in the information not allowed by that standard
- if the url is not included it should be because of the "intention"
- toolbars are subject to the same standards
I'm not disagreeing with you as much as just pointing out that I don't think EVERYONE agrees on these standards.As for ignoring robots.txt, that may not be the case. Conceivably you could get the url A->B link, save the metadata for both, signaling them as related, and then check both URLs against robots.txt to see if you should have them in the index. Then if url A is ".../search?q=torsorophy" Google's robots.txt disallows it from being indexed and only url B gets in but the link to "torsorophy" is still there from the metadata.
First just claim that the allegations are wrong (use 'Period.' and 'Full Stop.'). Don't try to explain!
Than redirect allegations as a personal attack (say it is 'insulting'). Don't try to explain why, either!
Refer to a document that requires paid membership, so only very few people if any can check (don't use a direct link). Gives you cheap credibility.
Claim that the experiment is fraud, although no one benefits financially. Don't back up the claim with data.
Counterclaim that your competitor is actually copying from you, ignoring the difference between ideas and data.
Now that the reader envisions you as the poor underdog that is insulted, is a fraud victim, and is copied, spread the doubt that Google is doing this because it starts to 'worry' (pose an open question like 'is it a coincidence'?).
End with the pose that your are not affected by this and concentrate on your business.
Never, ever answer the allegations with a simple explanation on why Bing shows bogus results because some Google employees search while the Bing Toolbar was active.
{Overlook that you admitted that your are prone to click fraud.}
It's like, "Son, did you steal those cookies?" "Mom, Justin always accuses me of stealing cookies, and I've been getting good grades lately, and last week Justin hit me for no reason, and don't you think he should be in trouble too?"
Answer the question, son.
Well, it was trained by having the Bing toolbar see people clicking those results on Google's search engine.
But I do think you're right that people are talking past each other and what outrages one person could very well be something that another simply doesn't care about.
Searching for "torsorophy" on Google brings up the Wikipedia page for Tarsorrhaphy because of Google's advanced spelling and error correction algorithms.
Searching for "torsorophy" on Bing will (or used to) bring up the Wikipedia page for Tarsorrhaphy because of Google's advanced spelling and error correction algorithms, because Bing watches how people use Google.
Two different kinds of innovation at play, one much more legitimate and impressive than the other. But it's not like Microsoft is in unfamiliar territory here.
When Bing adjusts the results as due to a spellcheck done by their engine, it notifies the user so they can correct it.
On your example, Bing search does exactly this: > We're including results for fast fourier transform. Do you want results only for fast forier twansform?
Both companies have good algorithms for spelling and error correction, but neither appears to be a superset of the other. (For example, Bing corrects "Kransas" to "Kansas", but Google does not).
I was addressing the fact that Bing has at its disposal some advanced techniques too. It's not just Google that does. And I illustrated this by pointing out that there are a lot of queries that do the spell correction on Bing that obviously weren't honeypotted -- as seemed to be implied by the original comment.
"Kransas" is Swedish for wreath, which is likely why Googles' spellcheck doesn't automatically correct it.
Bing however won't even show me proper results for Kransas even if I tell it "+Kransas".
One example of where they both 'fail' on an mispelling is "munday" which isn't autocorrected to Monday because munday is a proper noun existing in the english corpus---e.g. the City of Munday, the Munday family, etc.
Copying a leader isn’t as objectionable as denying it and I would argue that Google could split their search engine into independent entities—say: crawling, algorithm, interface. That way, competitors could try to outrun them and innovate on one or the other, separately: Firefox, Siri and many others on the interface; Cloudera or Amazon on the crawl, DuckDuckGo and physics labs on the algorithm (just spitballing there). What Bing did wasn’t that different from a grease-monkey script that would supplement bing.com results with Google‘s, it seems — or at least, they could have had the humility to present it that way.
Now, their pride made them confess to being easily gamed.
As I noted in a previous thread (http://news.ycombinator.com/item?id=2166256), Googler Amit Singhal's wording was a bit vague about whether URL trail data from things like the Google Toolbar, Analytics, Ads, or other systems ever affects rankings. (His careful wording, "put any results on Google’s results page", could mean merely that URL trail data never adds results to the set of all possible results. It could still affect rankings of pages found through other crawling.)
The use of such data is clearly allowed by Google's very broad privacy policy. I would love a clear answer from Google on this. It should be easy to give.
Robots.txt-sensitive crawling is irrelevant to this question, and whether their toolbar tracks clicks from other search engine result pages is only tangentially relevant, as one small example of the general idea.
And I'm not asking if they've ever done exactly the click inference Bing has done. Rather, I wonder if they're doing vaguely analogous indirect mining of revealed web user preferences via clicktrails. For example, noticing which sites were visited together or in certain order even without crawler-visible links between them. Or noticing which pages were viewed for the longest/shortest times. Or which pages seemed to 'end' a purposeful session. Or other deep-science stuff I can't even imagine.
I don't want them to reveal any proprietary secrets – just whether they have ever used (or would consider it legitimate to use) all the clicktrail data from all their many non-search tools to help with search quality.
Because I've long assumed that they do, and would be surprised if they didn't.
I would expect that they know this. The question is how susceptible does it make them? The clickstream is just one of something like a thousand factors that go into the ranking. Perhaps on any real page, it would take more fake clicks than a fraudster can generate to make a worthwhile difference.
Keep in mind that Google was experimenting with pages for which Bing had no information other than the clickstream, so the clickstream made a big difference there.
Erm, the Google test took pretty non-sensical keywords and put at #1 some completely unrelated website.
And then Bing copied this unrelated website result from Google.
The above quote from Bing (especially the bit in italics) is therefore a little confusing. Since, again, the entire point here is that Bing shown completely unrelated websites taking straight from Google. The results weren't relevant. Thus making the quote fairly contradictory in my opinion.
Yet another 'reply' by Bing, and yet again they seem to be trying to insult their way out of it via ad hominem and strawman arguments.
And yet again, Bing still haven't addressed the key point here: why Google results were taken and used by Bing.
This is how Bing's algorithm may have worked, and I wouldn't call it intentionally copying, flawed maybe and easily manipulated just like Google Bombing in the old days.
Though in the end, Google probably gets a net win cause they got "Microsoft copied us" headlines from some major sources with probably little follow up from the average reader. But the whole situation could have been handled better.
Keep in mind that we're talking about Google here. The Google of Google News, the web-corpus assisted translation engine, the Google book scanning project, wifi scanning, etc. Not that I'm negative on any of those projects, but based on the amazing degree to which their business piggybacks off the work of others it does make it difficult for me to take them seriously when they come out swinging at somebody else for piggybacking off of them, especially when the piggybacking going on here appears to be fairly minor and incidental.
Instead, Microsoft chose to attack Google 'personally' with non-substantiated claims.
Here's what appears to be the truth:
The Bing toolbar keeps track of what people click on after doing a search. Even when it is from sites that aren't Bing. If it is a search for something obscure (or made up) that Bing has no other data for, then that clickstream data can affect Bing results for that term. Most of the time, however, that is just one of thousands of inputs. It is useful data, but not a defining part of the Bing engine.
This was obviously a huge PR move on Google's part. The information was released to a massive search engine blog right before a big search engine event with both Google and Bing in attendance. Bing could have come out of this shooting straight and looking like the more mature party, instead the come across like they are unsure of themselves and have been backed into a corner.
Score one for the bullshit artists at Google.
Bing, "Powered by Google". It does have a nice ring to it.It can't be only that. The page where the clicks happen must be parsed for the click be linked to the query and reused at their search engine.
They transformed: "from google.com/q=X, people tend to click on example.com"
Into: "when searching for X, people tend to click on example.com"
It would be interesting experiment to estimate how heavily weighted the clicking is vs. the other 999 inputs. I would think a very obscure page would be fairly easily manipulated (Google basically showed that with their "honeypot"), but a more popular page with lots of non-zero entries in the other 999 inputs would be a lot harder to game.
The fact that only 7/100 of the honeypots were successful may mean that Google's motives may have been more complex and slightly less akin to righteous indignation - such as determining how and to what degree Bing handles "click fraud" type attacks. I just have a hard time believing that Google would turn 20 engineers loose on a task that can be so easily automated.
I disagree with describing this as an attack. Based on the kinds of search terms Google has said they used, they aren't attempting to make the Bing results any worse. There is simply no good search result for "hiybbprqag", so Google pointing to the Wiltern Seating Chart is just as good as pointing to nothing at all from the purposes of what any user who happened to Google hiybbprqag randomly would see.
At best, you could make the argument that by revealing that Bing is susceptible to these kinds of attacks, they are making it easier for others to attack Bing. It's certainly illegal to rob a bank, but what about walking down the street yelling, "The bank guards are all gone! The combination to the vault is 1-2-3-4-5!"
But it wasn't any user, it was a Google engineer and the circumstances were not ordinary. Those engineers were actively trying to get the results into Bing over the course of at least two weeks 12/17-12/31 according to searchenginland.com.
What I would classify as an attack is if, through a similar process, Google introduced bogus entries for real search terms, which they don't claim to have done.
IANAL, but if they have attacked Microsoft, that is probably a prudent legal strategy. Keep in mind that Google claims to have discovered this episode by noticing similarities between their results and Bing's which although it suggests and active program to monitor Bing, is hardly surprising.
On the other hand, however plausible it appears their claim about how they discovered it smells a bit of BS. It strains belief that Google never looked at the packets the Bing toolbar was sending home during browser compatibility testing. They lost their virginity a long time ago.
I wonder how smart the Bing Toolbar really is.
a) Search from Bing toolbar for hiybbprqag.
b) Bing does not display anything.
c) Bing toolbar remembers that term internally.
d) Go to www.google.com and type hiybbprqag.
e) Google displays one intentionally seeded link.
f) Google engineer click on that link.
g) Bing toolbar notices that shortly after user searched for hiybbprqag, s/he clicked on seeded link.
h) Bing toolbar sends that piece of that to mothership: There is relationship between hiybbprqag and seeded link.
...
i) After some time, Google engineer searches for hiybbprqag from Bing toolbar.
j) Bing looks up its index and there only one piece of evidence regarding term 'hiybbprqag'. It is not much, but it is all it has, so it presents it to the user.
Google accusations strongly imply (using words like 'stealing') that Bing simply scraps google.com for results, while reality is not so simple.
Now, bigger Microsoft problem is that it employs VPs like Yusuf, who cannot express simple facts and easily fall into corporate speak.
UPDATE: Here is good summary of what I wanted to say: http://directmatchmedia.com/google-proves-bing.php
"It was interesting to watch the level of protest and feigned outrage from Google. One wonders what brought them to a place where they would level these kinds of accusations." Feigned outrage... one wonders... If you really not understand this, you are not a hacker, you are not an engineer. Maybe it was your actions that prompted two valued Google engineers (Singhal and Cutts) to make these accusations?
"Before we explore that (so you gonna prove this point later on?), let me clear up a few things once and for all.
We do not copy results from any of our competitors. Period. Full stop." The end result is a copy of the results of a competitor. If you borrow it, steal it from your users, copy it from a guy in a raincoat with binoculars, observing Google results and jotting them down -- It doesn't matter! OK, let's say you don't copy results. Your index still contains copied results for a fact.
"We have some of the best minds in the world at work on search quality and relevance, and for a competitor to accuse any one of these people of such activity is just insulting." Bing copied Google's fake results. Your engineers either deliberately or accidentally made Bing copy those results. If you want less insulting: Google claims your minds are the best at copying others.
" We do look at anonymous click stream data as one of more than a thousand inputs into our ranking algorithm. We learn from our customers as they traverse the web, a common practice in helping to improve a wide array of online services. We have been clear about this for a couple of years (see Directions on Microsoft report, June 15, 2009). " That will make me trust your index a lot less. If engineers clicking manually on some results can make a difference in ranking, imagine what a botnet or dedicated spammer can do.
" Google engaged in a “honeypot” attack to trick Bing. In simple terms, Google’s “experiment” was rigged to manipulate Bing search results through a type of attack also known as “click fraud.” That’s right, the same type of attack employed by spammers on the web to trick consumers and produce bogus search results. What does all this cloak and dagger click fraud prove? Nothing anyone in the industry doesn’t already know. As we have said before and again in this post, we use click stream optionally provided by consumers in an anonymous fashion as one of 1,000 signals to try and determine whether a site might make sense to be in our index. " Ad hominems: Honeypot, trick, quotes:experiment, rigged, manipulate, attack, click fraud, attack, spammers, trick consumers, bogus, cloak dagger, click fraud, prove? nothing. If you need those words in a paragraph to describe how you got caught with your hand in the cookie jar, then you already lost.
" Now let’s move the conversation to what might really be going on behind the scenes. " Yes, move the conversation away from your duplicate results, the key issue at hand.
" Bing was launched nearly two years ago .to break new ground and help move the search industry in new directions. We have brought a number of things to market that we are very proud of -- our daily home page photos, infinite scroll in image search, great travel and shopping experiences, a new and more useful visual approach to search, and partnerships with key leaders like Facebook and Twitter. If you are keeping tabs, you will notice Google has “copied” a few of these. Whether they have done it well we leave to customers. But more importantly, we take no issue and are glad we could help move the industry to adopt some good ideas. " copies marketing paragraph. Rebuttal: Google didn't use its toolbar to "copy" your pictures and add those to their own background.
" At the same time, we have been making steady, quiet progress on core search relevance. In October 2010 we released a series of big, noticeable improvements to Bing’s relevance. So big and noticeable that we are told Google took notice and began to worry. Then a short time later, here come the honeypot attacks. Is the timing purely coincidence? Are industry discussions about search quality to be ignored? Is this simply a response to the fact that some people in the industry are beginning to ask whether Bing is as good or in some cases better than Google on core web relevance? " Bing is absolutely horrible for international searches. You serve up 40 different language Wikipedia pages for some terms. But you are saying, that Google deliberately timed your outings? That is like saying your brother saw you stealing cookies, but waited till he got a bad report card, before telling mom. And I thought a bruised academic ego was childish. So your rebutal amounts to: But mom! He got a bad report card! And that while you are struggling making your grades yourself.
" Clearly that’s a question that will continue in heated debate as long as there is a search industry. Here at Bing we will continue to focus on our customers, and try to provide some great innovation for consumers and the industry. " That's not the question and certainly not core to this case, which is your duplicate index. This is unprecedented in 10 years of search engine competition, and you want to make it about the timing? Let's make it about the 1000 ranking factors vs. the 200 ranking factor from Google. Lets make it about the million other results this could be happening for, but are impossible to prove with a "honeypot attack" and specific keyword? Lets make it about search relevance, how we are spoiled with relevant results, yet complain about the most relevant index on the web, one you use to calibrate long tail terms and spelling corrections?
This is pure nonsense, and so is the rest of your post. [edit: The business of] search is not about academia, or algorithms, or engineering, its about providing users with relevant responses to their queries. The best way currently to do that is through crowdsourcing--which is exactly what Google does through Pagerank. Now Bing took that a step further and creates an association between what a user searches for and what they end up finding relevant: This is exactly what we currently call "search". Whether that initial directory of sites came from Google or from some hand culled list is inconsequential. This is not "copying", this is doing exactly what they should be doing.
The point is Google doesn't own the link, or the form input, or the fact that many users clicked on that link after posting that term. The only thing they could be reasonably construed to "own" are the relative rankings themselves. Bing did not "copy" this.
People only click on the results that are there. Google put them there. People tend to click on results in roughly the order they appear on the page, which Google also determined. Getting the data through an indirect means does not insulate them from culpability.
Marketing and flashy features, two areas where bing has been investing a lot of money.
There are many sites on the internet that generate a set of links based on form data. Google is one of many in that respect. This technique is effective in gathering search information on this "deep web". Special-casing Google positively or negatively is the wrong approach here.
Btw good job with the disagree downvotes guys.
"[The business of] search is not about academia, or algorithms, or engineering, its about providing users with relevant responses to their queries"
And the food industry is not about the ingredients, the cooks and the recipes, but about providing users with a dish suitable to their tastes. In fact, we could do without the cooks, recipes and ingredients. er...?
"The best way currently to do that is through crowdsourcing". These are fine words to say that Bing uses rank on Google as 1 in 1000 ranking factors. Bing uses GoogleRank. Or CrowdRank in less polemic terms.
To me: it isn't about if Microsoft is "copying" or not anymore. You can win that semantic game. What you can't win or explain away is the end result: the duplicate results and grammar corrections. The process by which this happened (Microsoft can hack my microphone and record keystrokes) is irrelevant. Bing uses rank on Google as a ranking factor and Bing contains duplicate results.
I apologize for the unnecessarily mean-spirited reply, it was uncalled for.
>And the food industry is not about the ingredients, the cooks and the recipes, but about providing users with a dish suitable to their tastes. In fact, we could do without the cooks, recipes and ingredients. er...?
You're wrongly combining "food" and "the food industry". Food is about the end result, regardless of how it came about. The food industry is about all those things you mentioned. The same goes for search.
>it isn't about if Microsoft is "copying" or not anymore. You can win that semantic game.
You're right about this, "copying" or not is purely a semantic game. The problem I see is that Google started it intentionally. They could have used more precise terms and actually started a conversation about the real issue here, which you correctly identified, as Google results showing up in Bing. But they chose to go the sensational route.
This is my reply to moultano that addresses your point:
The value of a search engine isn't any particular result, or any set of results. Its the quality of all the results over time. If Microsoft's algorithm picks up a tiny amount of signal (ahem 1 of a 1000) indirectly from Google's results, this does nothing to artificially inflate their position off of Google's back. There's nothing inherently wrong about using user signal for this.
There are many sites on the internet that generate a set of links based on form data. Google is one of many in that respect. This technique is effective in gathering search information on this "deep web". Special-casing Google positively or negatively is the wrong approach here.
...have nothing to do with the logical merits of the opponent's arguments or assertions... The proof was concocted by these engineers.
...but can also involve pointing out factual but ostensible character flaws or actions which are irrelevant to the opponent's argument...
The way his statement could be read is: Matt Cutts and Singhal are no more than spammers, tricking Microsoft with cloak and dagger tactics that prove nothing. Which constitutes a clear Ad Hominem.
(Also, it's spelled Singhal, not Singhai)
This was the phrase that almost made me want to do another round of posts. As the person on the panel from Google, I can assure you: it was not feigned outrage. It was real frustration.
Another thing that did not occur is click fraud. That term already means something, it involves pay-per-click advertising, and it doesn't have much to do with Google's experiment.
Don't let this guy distort the language we use to discuss the issue.
An alternative is that the bing toolbar is collecting 2-tuples "<search_string_in_toolbar, next_href_clicked>" and sending these back to microsoft (regardless of the search provider selected). I would consider this "click stream data" and seems to agree with statements in the above article. Additionally, since it doesn't involve parsing hrefs it seems like the easier solution (and the one I'm going to tentatively assume by Occam's razor).
Just to be clear, it does not have to be the case that the bing toolbar is collecting data entered directly into the google search box. Indeed, google could have tested for this by having a control group where they entered search queries without using the toolbar search box, however from the details released so far we cannot conclude they did this (read: google hasn't released enough specifics about the test they conducted, and it's far from being reproducible with the current details).
Additionally, why do many in the HN community think this google specific? I understand that google isn't claiming "bing copied just google" but that seems to be the consensus within HN and the arguments of foul play (see the countless posts asking if there exists code that specifies google: "if string contains 'google' then..."). I'd like to see a test where users entered a pathological string in the toolbar with an alternative search provider specified (EDIT: not bing or google), clicked on a low ranking result (low ranking, or not appearing at all on bing for those keywords), and see if it pops up higher on bing at a later time.
Google seemed to think so, but even if Microsoft was treating the data they received as neutrally as possible, it doesn't change the request made of them. Google wants an exemption.
BING!! uses GoogleRank as a ranking factor. Perhaps they even adjust the weights of this factor, according to what the press/blogosphere is writing about Google's index quality...
ps: FYI, Google does record the clickstream within their own search results.
The only great thing I see is that the blog post actually confirms what Google said. Because otherwise, it would have been easy for Microsoft to put the finger at the point where Google is wrong.