Facebook employee responds on robots.txt controversy
petewarden.typepad.com
petewarden.typepad.com
There are a couple of things I want to clarify. First, we genuinely support data portability: we want users to be able to use their data in other applications without restriction. Our new data policies, which we deployed at f8, clearly reflect this (http://developers.facebook.com/policy/):
"Users give you their basic account information when they connect with your application. For all other data, you must obtain explicit consent from the user who provided the data to us before using it for any purpose other than displaying it back to the user on your application."
Basically, users have complete control over their data, and as long as user gives an application explicit consent, Facebook doesn't get in the way of the user using their data in your applications beyond basic protections like selling data to ad networks and other sleazy data collectors.Crawling is a bit of special case. We have a privacy control enabling users to decide whether they want their profile page to show up in search engines. Many of the other "crawlers" don't really meet user expectations. As Blake mentioned in his response on Pete's blog post, some sleazy crawlers simply aggregate user data en masse and then sell it, which we view as a threat to user privacy.
Pete's post did bring up some real issues with the way we were handling things. In particular, I think it was bad for us to stray from Internet standards and conventions by having an robots.txt that was open and a separate agreement with additional restrictions. This was just a lapse of judgment.
We are updating our robots.txt to explicitly allow the crawlers of search engines that we currently allow to index Facebook content and disallow all other crawlers. We will whitelist crawlers when legitimate companies contact us who want to crawl us (presumably search engines). For other purposes, we really want people using our API because it has explicit controls around privacy and has important additional requirements that we feel are important when a company is using users' data from Facebook (e.g., we require that you have a privacy policy and offer users the ability to delete their data from your service).
This robots.txt change should be deployed today. The change will make our robots.txt abide by conventions and standards, which I think is the main legitimate complaint in Pete's post.
I just wish we could have had this conversation a few months ago, before you guys threw your lawyers at me.
You don't have an "agreement" at all. An agreement requires that two parties actually, um, agree. You've published a statement where you assert certain rights and imply that you will sue anyone who accesses data on your site in a way you don't like. You may get away with that, regardless of the legal merits of your position, because you have more money for lawyers than most people you're likely to sue. But don't try to dignify what you're doing by calling it an "agreement". It's like an extortionist telling me that we have an agreement that he won't break my windows if I pay him protection money.
I'm a little concerned about how we (80legs) fit into Facebook's permission form, though. We're not using or selling the data ourselves, but we can sell access to an already-setup web crawl - so how does that work? We're like a special case of a special case. Never an enviable spot to be in.
IP, location, host of the site you are viewing whois details for the site caching rules robots txt rules for the page page rank blah, etc
Probably as a dropdown info panel.
I'm surprised if this doesn't already exist, though.
For example, our ChangeDetection.com site has a couple hundred users monitoring pages on facebook. (User-Agent: ChangeDetection). We always honor robots.txt so if we are not whitelisted all these monitors will be disabled.
Edit: looks like you already have posted contact info in robots.txt
Based on my interaction with website API developers, most of them honestly believe that they are building an open web. But let's face it, the API is a benefit to them and creates a critical dependency for the user of the API. It's basically vendor lock-in.
I'm not sure what you want Facebook to change.
Speaking as someone who's working on leveraging the Facebook API in a commercial product, this leaves me feeling like I'm opening myself to a lot of legal exposure if Facebook subjectively decides that my service poses even a minor threat to them. Given that I'm bootstrapping, there's no way I'd be able to put up any sort of legal fight what so ever against a company as well-funded as Facebook.
Though if you've made it clear that only x, y and z can crawl your site, and someone spoofs, say, y, then it would be easy to demonstrate that someone has done something they know they shouldn't.
and not only can the bot lie, it can disregard the robots.txt file altogether. just like the terms of service document for humans, you can choose to disregard it & deal w/ the consequences (blocked IP's, lawsuit, etc).
robots.txt is just a version of the TOS that computers can read.
I think Facebook's ToS is an appropriate way for them to send messages about what they'll do if you make certain uses of the content you retrieve from their site. However, these ToS documents aren't omnipotent: they can't restrict some fair uses to which you might put the data, for instance, or bind you to silly terms. IANAL, but I think the ToS is probably best understood as an intent to use still other sets of rules (perhaps selectively) if you do certain things with retrieved data. Disagreeing with that policy is certainly possible, but I don't think robots.txt has much to do with it.
Facebook's policies aside, it's interesting to note that this criticism was written in response to a letter from Blake Ross- the founder of Firefox.
this new model that facebook is trying to push isn't scalable. it favors the big guys and it's bad for the open web.
you shouldn't have a TOS that contradicts your robots.txt. period.
Of course they tune their page so that search crawlers can best index the information, so does everyone else on the web.
However, they also provide an API, with clearly defined terms of use, which you may use to get information. Basically your complaint boils down to you don't want to use the methods they've set up for you to access their data, and you're complaining about it.
As for comments about fair, true spirit of the internet what have you... I don't think the true spirit of the internet has ever been everyone has to give everything away in every possible imaginable way. And, in the end, it is Facebook's data, they make that clear before you ever start adding data to their servers. So does just about everyone else who allows you to submit data to their servers.
This shouldn't be confused with how difficult it is to remove data and accounts from their system, which is a giant pain in the butt, or the fact that they've made drastic changes to the public nature of the data after keeping it private for so long. That's all just a giant mess.
User-agent: *
Disallow: /ac.php
Disallow: /ae.php
Disallow: /album.php
Disallow: /ap.php
Disallow: /feeds/
Disallow: /o.php
Disallow: /p.php
Disallow: /photo_comments.php
Disallow: /photo_search.php
Disallow: /photos.phpSuing people because you're too lazy to only give your content to people who you want to receive it is wrong.
That's not to say their C&D is valid or enforceable. It just enters into another legal bin that's a bit more hairy.
Unfortunately, the legal precedent for all this is murky at best. Past cases have been very specific to the details of those cases. No line has ever been drawn about what is ok and what is not.
Exactly. It will be interesting to see where things end up playing out in the "leaving yourself wide open" department.
It's very much parallel with the "using open wifi points" debate.
But are those terms an enforceable contract? Probably not. Merely making available a document that purports to be a contract does not make it so.
Not really true, their policy clearly says:
"You own all of the content and information you post on Facebook"
But clearly from Apple's example it is possible to build a profitable product on a closed system.
Admittedly it doesn't seem "fair" or true to the spirit of the internet. but it shouldn't be any surprise that a company is going to do what is best for that company forsaking all others.
Change, in general, doesn't come about by people sitting around doing nothing but saying, 'the market will decide.' If you take the time to become an advocate for a certain 'side,' you can affect the market through public perception.
[aside: People are not always logical agents with access to full information like most economic theories seem to assume. To me this comes off like a lot of physics where things are assumed to be on a 'perfectly frictionless surface,' 'in a vaccum', etc.]