Nate Silver: How Much Does Bing Borrow From Google?
fivethirtyeight.blogs.nytimes.com
fivethirtyeight.blogs.nytimes.com
With much writing, he essentially says the following:
1. Microsoft's defense contains no information because they don't say how they weight the data learned from Google searches.
2. It's hard to say how Microsoft is actually using the data learned from Google searches.
For the crowd here this is essentially common-sense.
(edit: in point two I changed 'impossible' to 'hard')
In many cases where they have no other data, they clearly weight Google's single data point high enough to create a whole SERP around it. The claim that 'this is just one input among many' implies that some kind of cross-referencing or corroboration takes place. In this experiment that obviously didn't happen.
This really doesn't tell us anything about how highly it is weighted relative to other data though, since obviously there will be no other data to weigh in on a nonsense search term.
Whether they throw up new SERPs in reaction to data from other single sources might be interesting. But if true, it's also not very complimentary.
However hand they issued twenty engineers laptops which is more consistent with an ongoing effort than a single afternoon of Bing and beer - particularly when you consider that if there was a single method applied and it was known beforehand, one engineer could have automated the whole exercise.
The inability to accurately measure honeypot injections is a bit odd. As pure speculation, it may be that they were not sure if the methods they used to get Bing to show honeypots number 8 and 9 were legitimate.
It's common sense that they are using data from Google. That has pretty much been proven conclusively. Nate Silver telling us this again is not interesting. The possibility that he might try was.
I actually think Silver makes a useful contribution, mainly because of his background and gift for illustrating data perception problems. If you're looking for hard facts and (more) smoking guns then you would naturally be disappointed.
Though I admit I may only read the first couple hundred and skim about halfway through before closing the tab.
In the case of the honeypot pages, clearly the number is 100%.
In the cases of other results pages, we merely have Google's claim that the similarity has been rising over time, and that they attribute it to info extracted from Google.
Using information that people provide about where they click (hopefully with some kind of informed consent), in order to improve your algorithm, seems reasonable enough. If a large part of the info ended up coming from Google's algorithm, it starts to seems sketchy, and it's legitimate for Google to whinge.
Not that whinging ever stopped any other company, particularly Microsoft, from trying to take advantage of other companies' IP.
Well...
'over the next few months we noticed that URLs from
Google search results would later appear in Bing with
increasing frequency for all kinds of queries:
*popular queries*, rare or unusual queries and
misspelled queries.'
-- http://goo.gl/Bi0JH [googleblog]
Now it's possible that Bing really did have no data for these 'popular queries' but I don't think that's an argument anyone would like to make. The alternative interpretation is that top-ranked pages in Google get added prominently to Bing after some interaction involving the toolbar, even when there are lots of other possible matches.You can still say that this would represent just a single data-point among many but unless you think Google is lying, you can't say that it only happens with rare searches.
BTW: I'm not arguing that there is direct evidence that the mentioned effect on popular queries is related to the toolbar. But the effect is there.
The should search for Justin Bieber and add in the Bieber search results some really horrible link.
MS should have a lot of relevant info for Justin Bieber, so the Google data should be rated lowly. If this non-sensical result makes the first page then you can begin to assume that their weighing it heavily. But when your query is "erftqnvpwedf" -- that just means that the only relevant info is coming from the toolbar. Bing would show that result even if the toolbar only accounted for 1/(2^50) of the total relevance.
I suspect Google did this test and has nothing to report.
This is pretty straightforward and I'd be surprised if Google didn't have the infrastructure in pface to do this today (like literally today).
Also do things like 'radix-2 fft', something where there is a lot of additional signals (I suspect), yet something that other toolbar members probably aren't searching for. So the toolbar data MS gets is strongly skewed to your results.
This is the very thing that Google is saying that Bing is stealing with their clickstream data, and it's also the one that you're agreeing would occur - "that the only relevant info is coming from the toolbar" for highly unlikely queries.
So it seems to me, looking at your argument, that you actually agree with Google's point.
Google is trying to argue a huge PR point by saying, "Bing copies" with the ramification being that when you search on Bing you're really searching on Google. If Google said this, "For extremely rare searches where Bing has few, if any, good signals, clickthroughs from Google will be weighed in such a way that these results may make the first page".
I'd buy that 100%. The Bing team may even buy it. Google seems to want to start with that thesis, but then try to shove the whole camel in too.
For example, lots of SEO and normal research studies the top Google results, or keywords revealed from referrer info and other logs, and then mimics those results elsewhere on the web, including outlinking to top results found via Google.
So anything prominent in Google's results will be echoed elsewhere in crawlable form, and probably associated with the same terms by which it was found at Google. Then those once-unique results will 'later appear' in other resources.
You can't be the dominating giant of the industry – the "single point labeled G connected to 10 billion destination pages" (as Blekko's Skrenta calls Google) – and not have this leakage happen. It goes with the territory.
Maybe because it's a good result? Possibly Bing arrived at it the same way Google did? Or has Google been slipping paper streets into their queries for longer than they're telling us?
In short: how did they know that it was "URLs from Google search results" that they were seeing in the Bing results?
If it was a popular query, and both google and bing made the same "mistake", doesn't that seem a tad suspicious to you?
Later, someone comes up to you and asks which is the best italian restaurant on that street. You recall that the couple with the restaurant guide picked Batali's, and so you point them to it.
Google would say you are ripping off the restaurant guide. Microsoft would say you are just using the observed behavior of the first couple to guide you.
That is, if any learning system is observing click-stream behavior from users and mining it for relevance evidence, I'd expect it to ultimately home in on the true weight of each piece of evidence in that click-stream data. Since Google's contributions to that data are likely to be highly relevant, any good machine learning system is going to learn that they are relevant and start recommending them.
In effect, then, what Google is arguing is that if Bing's machine-learning algorithms are correctly inferring that results that happen to come from Google are highly relevant, Bing should blind itself to that knowledge.
I'm not sure that's good for anyone except Google.
(For a completely uninformed guess about why Google might be interested in raising copycat claims about Bing, I go out on a limb here: http://blog.moertel.com/articles/2011/02/02/the-google-micro...)
Or to put it another way: If google did not exist do you still have a (good) search program? If the answer is no, then there is no reason for them to exist.
No, if your learning system is working properly you should be using Google's data only to the extent it is legitimately observable and more relevant than anything else you're feeding your system. And, for lots of searches, Google's data leave much room for other sources to be more relevant. In the limiting case, when you're feeding your system everything that Google is feeding its, you should almost never return Google-derived knowledge because your system should almost always be able to come up with greater relevance from knowledge derived from primary sources.
Your final question is on the right track, but I'd suggest a small tweak: If Google didn't exist, would the system still offer highly relevant results and, if Google did exist, would the system be able to learn from Google-supplied knowledge, to the extent allowed by law and terms of service and so forth, to offer results at least as relevant?
Microsoft has lifted Google's SERPs verbatim for a particular search ... what else is there to discuss?
The trivial addition to their experiment would be to add those nonsense words to a few Wikipedia pages, and click the links with the Bing toolbar installed, and see if those pages show up in the Bing results.