A surprising opinion on the ‘Bing is copying Google’ controversy
fury.com
fury.com
Hasn't this been at the core of the controversy all along?
Counterargument:
It's user action on a page created by google. Say I had nothing but this user data - and say I had enough of a sample. I could build a pretty damn good search engine out of that. But I know nothing about how to provide search results. I haven't even got a crawler. Yet here I am with a search engine.
edit (to finish the follow through):
If I can with this method build something about which I have no actual knowledge - then it proves the method is piggy backing on someone else's hard work. And THAT is what matters. It doesn't matter if it is only one of many signals - it doesn't matter if it's not copying Google's algo directly. What matters is that it is still relying on someone else's knowledge to improve Bing's product. That's shitty.
This argument clinches it for me.
The fact that Bing can't even acknowledge the opposing side - given that at the very least intuitions are going both ways here - is a clear sign of bad faith to me.
What's skeevy in the MS case is that the relied-upon-knowledge and the product are so similar...but still, Google's uber-umbrage strikes me as odd and tone-deaf, especially given its dominant position in the market.
I'm not sure there is any way for Google to Tell MSFT to not use their search results in the Bing search engine.
If Microsoft had just been a bit more upfront with the fact that they were using IE user's google clicking behavior to improve their search engine, I suspect there would have been much less furor.
But it seems very hypocritical to get butt-hurt on behalf of Bing toolbar users who are having their movements tracked.
(BTW, I don't find Google's observance of robots.txt to be particularly compelling or telling, because they've never been in a position where ignoring it would be significantly beneficial to them, as far as I know.)
I agree with you though, if I, as an IE9 user wish to submit my click track results to MSFT for analysis so they can improve search results - that's fair game.
But, It's not clear to me that MSFT should be able to review what the user was searching on before they clicked on that data. Now they are actually using Google's search Data + the user's click traffic. I think they cross a line there, particularly if they aren't willing to come clean and admit that's what they are doing, and make it clear that they are sending your Google Search queries + your click traffic back to Redmond.
How many people on HN were aware that Microsoft was doing that with IE? Click Traffic, sure - But I didn't know they were sending my Google Search queries back to HQ.
It's voluntary, you can opt out of the anonymous data reporting. The data they are using belongs to the user, so it's the user that can opt out, not Google.
Google could send a DMCA Takedown or other Cease-and-Desist letter.
The intuition is that if a method allows you to build a search engine without actually knowing anything (besides the method itself) - then that method is piggy backing off someone else's work.
Yes - google piggy backs off other people's work - i.e. linking to other sites to indicate quality. But this is not piggy backing off a search engine. It's not piggy backing off the technical work that someone else developed for a search engine product. So it's not a counter example to my argument.
I don't care about piggy-backing or fairness; I care about maximum public good.
So here's the situation. Google wants Bing to stop harvesting its results. Why? Very broadly, there are two scenarios:
- Bing is adding value - Bing is not adding value
In the latter scenario, Google should not care. No one will use Bing, and it's a non-issue.
In the former scenario, it's not clear at all to me that Google should have a monopoly on the information entered and sites visited of people that happen to visit their site. Google is trying to assert ownership, in a fashion, of the link between the users query and the site they visited, because the site they visited happened to be shown to them by Google. I see no reason to grant them this right.
The fact that Google acts outraged be Bing's behavior strikes me as particularly rich, given that Google itself is a notorious and cavalier harvester of data others would consider "theirs".
(Edit: This is pure, uninformed speculation, but I wonder if Google's (to me) odd outrage is because they place such a premium on algorithms over human-generated content. I.E., the idea is it's fine to harvest other people's human-generated work ("information wants to be free"-style), but harvesting other's algorithm-generated content is verboten, because that's essentially stealing the algorithm. I have no inside knowledge, but this would agree w/ the pop-culture characterizations of Google.)
Bing is not adding value with this method - they are thieving value. By thieving it they reduce the reward for the value that google provides. As such they reduce the incentive to provide genuine innovation.
Having said that - I don't necessarily disagree with the view that Google's position is a bit rich given the data that they do harvest from us folk without remunerating us for that work. But that doesn't make the argument against Bing any weaker. Two wrongs n rights n all that.
When it comes to PageRank, the implicit offer is: Allow google to use your links (intra-site and extra-site) in its algorithm, and it will make your site searchable via Google.
If you don't want Google to use your 'hard work' of collating and vetting links to other sites, then disallow the Googlebot via robots.txt
On the other hand, Microsoft refuses to allow anyway to disallow click traffic patterns involving your site to be used in its algorithm. They are thus mining the links from your site to another without even having to renumerate you by giving you a chance to be indexed in their engine.
Because that information doesn't belong to the site to which it is associated, but to the user.
A webpage author links to another as an act of conscious recommendation. The one who added the link added it for the sole purpose of others to make use of it. It would be a stretch to claim that Google serves search results for its competitor to make use of.
Next is the issue that Bing is not scraping Google but using user clicks. But here users are just a means to an end of scraping. That someone is doing an act by way of third parties does not take away from the fact that one is still doing it.
I am all for having someone give Google the run for its money, but I want that to be driven by genuine technological innovation. Not by being a El-Cheapo knockoff of the market leader. Some search engines are beating Google in niche markets by being better than Google, which is excellent.
I dont think, Bing re-serving Google results, brings any innovative pressures to the market. Especially when you know that whatever innovations one brings it will get replicated by piggy backing. I wouldnt want to be in a business where this is true.
What worries me is that the main players will start engaging more in how to inconvenience each other rather than building better products. I have been fearful of the fact that Microsoft would someday tweak IE so that Google does not work well on it. They have done this for a few other sites but haven't done it to Google. I think initially in fear of starting a race, but now with other browsers beginning to rule the roost thats moot.
During that time Google has managed to innovate fast enough and implement well enough to grow into a multi-billion dollar corporation.
The whole point of generating 100 unlikely search terms is the fact that these wouldn't be items that exist in any web page on the web. Hence if one searches for them it should not return any results!
How is monitoring people's clicks on google's search results not copying ?
On a different note, I am of the opinion that this practice is acceptable, only as long as one acknowledges it. The fact that M$ is not, to me is appalling.
But people do have a bias to click on the first result because it's the first result. From http://www.useit.com/alertbox/defaults.html :
> 42% of users clicked the top search hit, and 8% of users clicked the second hit. ...
> [When the researchers] swapped the order of the top two search hits ... users still clicked on the top entry 34% of the time and on the second hit 12% of the time.
I didn't read the paper (just read the abstract, I know..), could anyone explain what are they doing so bad in that? Thanks.
Bing/MS can choose to dishonor that convention, either directly or indirectly as they are now with the Bing toolbar, but I think that's a mistake on their part. Google (and Bing.com for that matter) has used the most widely accepted mechanism, robots.txt, for ensuring that content it does not believe should be indexed is not. It abides by that convention with other sites and expects other search engines to abide by that convention with respect to google.com in return.
From here the following things can happen:
1. Microsoft agrees that robots.txt should govern clickstream data from the Bing toolbar. The world returns to sanity once again.
2. The search engine community (MS, Google, etc.) agree on a new standard similar to robots.txt specifically for governing the use of clickstream data, sites are updated with new directives allowing/disallowing such use and the world returns to sanity.
3. Microsoft specifically denies that they should be limited in using clickstream data from any source, through convention or through any means. Search companies and other sites are forced to fall back on other means to achieve the same results and things get messy (for example, google blocks all users who have the Bing toolbar installed, they prevent the Bing toolbar from being installed in Chrome, MS retaliates, etc.)
For the life of me I cannot think of a sane reason why MS would choose option number 3 other than that they are insane, horribly myopic, or just plain dumb.
I can imagine a reverse sting operation Bing could run: 1) Set up similar faked honeypot pages that only Bing knows about. 2) Send emails to Gmail accounts owned by the Bing management team. 3) Click on a few of those links. 4) Watch as Google mines the clicks and ends up with the links in their search index. 5) Sit back and accuse Google of reading private emails from Bing management to improve their search results.
It's exactly the same thing, the data is owned by the user, and in both cases the user will have agreed to a terms and conditions that allows the companies to collect stats from that data.
If you want to control what a search engine does with a page, you should use <meta name="ROBOTS"> and rel=nofollow, neither of which Google includes in its results pages. It would be reasonable for Bing to look for and honor these directives when collecting their clickstream data, and Google isn't availing itself of the option.
robots.txt also blocks indexing of non-html files.
The point is: up until now search indexing has had a robust opt-out mechanism. Bing has changed that and has muddied the waters with vague hand-waving. The question still remains: should there still be a way for sites to opt-out of search indexing? If so, then why can't robots.txt continue to be that mechanism? If robots.txt can't be that mechanism then what should be the new standard. If not, then that opens an entirely new can of worms. A can that I think we're better off not opening.
Fox also suggests there could be a robots.txt-like standard where sites declare they want to opt their users' activity out of any such analysis. That strikes me as a bad idea: users ought to own their own interaction trails.