How I got sued by Facebook
petewarden.typepad.com
petewarden.typepad.com
This is a very real problem with law surrounding emerging business practices, esp here in the US.
Ultimately you can only pioneer whatever you can afford to defend in court. Facebook, or any other BigCo for that matter, can assert that you can't do X and it's up to you to fight it in court... Even if there is prior behavior such as with this case. Clearly Google does the very same job and doesn't have an agreement with Facebook to spider their site.
But if you can't afford to defend it and bring up the prior examples in a court, then Facebook - or anyone else for that matter - can stop you.
This is ridiculous. Can someone release the data so this can be tested in court? EFF?
I've never used facebook, but Section 9.2.6 of their terms says You will delete all data you received from Facebook if we disable your application or ask you to do so.
I don't see how that could be encoded in robots.txt.
Either facebook wants to allow people to crawl it, or they don't. robots.txt should be binary - yes or no.
Statement #2 is a false dichotomy.
The Terms for facebook is a terrible document. It is written to be intelligible to humans but is full of ambiguity and undefined terms. Any lawsuit about the theory "robots.txt did not forbid my actions; therefore they are legal" would probably disintegrate into how the Terms language is interpreted.
Their lawyers probably leapt from high windows when it was released.
For a start though, you could easily state that facebook freely allow access to data, without requiring you to read terms&conditions.
You could argue that by freely allowing access to all of their data, without requiring you to read and agree to the terms of usage, then their claims have no basis.
If they did want to restrict access, or make sure every crawler had first agreed to T&C, it wouldn't take long for them to add that.
It hasn't been tested in court, and there's a truly excellent chance Facebook would lose - either in terms of the court or in terms of what's left of their privacy reputation - but that doesn't matter one little bit. They have attorneys and money, and that's all that matters in this instance.
Really. The facile assumption that "it's possible to aggregate thus it's OK to aggregate" is exactly the way normal people think (by which I mean, tongue-in-cheek, us), but corporate attorneys see all this in terms of power relationships and contracts. As far as they're concerned, the poster took pictures through their front windows, and they're damn well going to threaten him with kneecapping until he gives them his negatives.
Check out this site (just found it via search): http://www.canyoucopyrightatweet.com/
Disallow: /directory/people/*
Disallow: /directory/pages/*
Done. Google crawls Facebook. Every social media monitoring company crawls Facebook (and every other social network). Tons of other companies do the same every day. I know this because many of these companies are our customers.
What really irks me is that Pete collected valuable data and whoever he shared it with was probably able to derive added value from that data - in ways that Facebook is not doing. Facebook is prohibiting value creation.
Facebook doesn't give a tinker's damn about prohibiting your value creation.
"If you collect information from users, you will: obtain their consent, make it clear you (and not Facebook) are the one collecting their information, and post a privacy policy explaining what information you collect and how you will use it."
"You will only use the data you receive for your application, and will only use it in connection with Facebook."
"By "application" we mean any application or website (including Connect sites) that uses or accesses Platform, as well as anything else that receives data." - note that by this definition a Facebook crawler is an applicatiom
"You will only use the data you receive for your application, and will only use it in connection with Facebook."
"You will have a privacy policy or otherwise make it clear to users what user data you are going to use and how you will use, display, or share that data."
"You will not transfer the data you receive from us (or enable that data to be transferred) without our prior consent."
"You will make it easy for users to remove or disconnect from your application."
etc.
These things are easy to look up, it took me about five minutes to find them. In fact, they have the right to limit the way even publicly faced content on their site is used. Things like copyright and data protection laws come in mind. You might not agree with that but it is absolutely in their right to do so.
The FB Terms & Conditions don't mention robots.txt at all.
If you run a public website, you have to accept that people may crawl your website. If you want to prevent it, don't let them crawl it. That's what robots.txt is for.
Everything from stock prices to the temperature is in theory uncopyrightable, but you can still run into trouble if you just scrape some site and republish it.
This could have been an interesting case.
And regarding robots.txt - it's by no means anything more than a suggestion and common courtesy.. it's a way to suggest to crawlers what to ignore and how to behave with your site, to their benefit and yours - it's not by itself a legally binding, well, anything..... it was just a convention for all parties to unofficially cooperate.
It depends, re: the oft cited example of phone books being uncopyrightable,
If I buy a map, can I use it to be a guide, giving people directions and make a living?
If I buy a story book, can I use it to read to children and collect money from their parents?
Or, if I spend some time to index all the books in the library, can I sell the index and make money? Or did I violate the rights of either the library or the books' copyright owners?
I think crawling and making use of the crawled data is not offending the copyright law because it's not actually a copy of the data. Rather, it's a service to transform the data and help other people get to the original data in an easier way.
Technically, they can crawl the data on the spot when receiving a request from their clients. In order to speed up the process, they somehow cache the data crawled from earlier occasions. But that's technical details, which is encapsulated. Otherwise, they could have stored URLs and offsets in the pages, rather than the data itself. I don't think that breaks the law, or there will be no way to avoid breaking the law to refer to anything.
Another possible related issue is the question of derivative works (e.g. the index of a library). There's a Wikipedia article about derivative works, it's a bit too complex to summarize here: http://en.wikipedia.org/wiki/Derivative_work
Both films and music (plus songwriting too) have been explicitly enshrined by US copyright law as having separate rights for home use and public performance. Each has multiple spheres of statutory licensing organizations with special exemptions from anti-trust law.
I'm not very familiar with US copyright legislation but I assumed that the copyright laws there are similar without checking any sources - which was a mistake apparently.
But, really, they can't stop it. An when (not if) nominative aggregated data is out, maybe the general population will begin to actually understand how expensive Facebook actually is.
Well, they claim it is their right. The courts will decide if it actually is.
Plenty of places make claims to rights and make you sign waivers (dojos, for example), but fail in court. This could be interesting.
It's probable they would lose. But it's far, far more probable that Pete would be bankrupt well in advance of that point in time, and that's how the court system works.
"Their contention was robots.txt had no legal force and they could sue anyone for accessing their site ..."
Lone data collector dude:
<blink> <blink>
It turns out the lawyers are right. Huh.
This was back in the day when everybody was using the unix network (pine for email, etc). Many people had webpages on their university accounts and would also store homework online. A very popular thing to do was to go into someone else's home directory and copy their .fvwmrc or .bashrc file (yes this was a few years ago).
The IT people came up with a new policy that said you were not allowed to look into other people's directories without explicit (possibly written) permission from the owner. They said that copying files could be considered "cheating" and possibly copyright infringement and you'd be referred to the university disciplinary court. I assume this all stemmed from fear of cheating.
In any event, I got into a long debate with the IT people about the whole point of unix file permissions. Basically: if you don't want people to look at your stuff, don't let them. As you might guess, when dealing with university IT, I was on the losing end of the argument.
From my point of view, there has been lots of individually useful rules that have been added to the court system, that when taken in aggregate have negative effects. There are lots of pre-trial motions that have specific uses, designed to stop a problem, or plug a loophole in the law. The problem comes when you have multiple dozens of them. Or even just a few big ones.
For instance, discovery. The idea is that fairness is not helped by surprise evidence. So, the solution is that both sides get access to the other side's evidence. But how do you do that? You have lawyer time, calendar time, etc to collect and go through it all.
That's just one (although one of the larger ones). Add in requests for everything else, and you can see how much lawyer time is "wasted".
They would state that Google would be happy to reindex Facebook, but it would take several weeks for their lawyers to meet with Facebook's counsel, draft documents, reindex privately, audit their cache to ensure compliance, etc.
We could then sit back and watch how fast Facebook backpedals.
The law acknowledges standard practice and expectations. That's what it is built on.
If you put something up on the web, then although you have copyright, you're also implying permission for people to be able to do what they normally do with it (at a minimum, view it in a web browser).
Similarly, if you've put up a robots.txt, then you're also quite clearly giving permission to crawlers (within certain bounds). There's also reasonableness and taking into account what robots.txt is capable of, of course. You clearly aren't implying permission for me to DoS you.
Explicitly agreed terms trumps this. If it could be shown that the person operating the crawler had agreed to an AUP, then that would change things. I imagine that this would make this case quite different from Google crawling your site.
Read about Zuck's wget magic here, and how he illegally scraped (guarantee it was a TOS violation) all the Harvard online facebooks in 2003: http://www.scribd.com/mobile/documents/538697
As Balzac said, behind every great fortune lies a great crime :)
If it's real (it's a pretty apalling read into the guy's mind), what a great example of why you shouldn't keep all your private thoughts online...
Claims: to remove all the content indexed by Google's crawlers on the newspaper's websites."
http://www.infoniac.com/offbeat-news/google-list-of-class-ac...
Google's response: "Of course, if publishers don’t want their websites to appear in search results (most do) the robots.txt standard (something that webmasters understand) enables them to prevent automatically the indexing of their content. It's nearly universally accepted and honoured by all reputable search engines."
http://googleblog.blogspot.com/2006/09/about-google-news-cas...
May be you can get Google to file an amicus brief...
"Outcome: Google had to remove the plaintiff's newspaper content from its database within 10 days or face fines of 1,000,000 Euro per day. Google had to publish "in a visible and clear manner and without any commentary from her part the entire intervening judgment on the home pages of google.be and of news.google.be for a continuous period of 5 days within 10 days... under penalty of a daily fine of 500,000 Euro per day of delay". Google had was awarded the costs of the expenses of 941.63 Euro (summons) and 121.47 Euro (costs of thy proceedings)."
I'm no Facebook hater, I actually use Facebook often but when they pull these kind of stunts it really upsets me. I would imagine it would upset most freedom lovers on this site as well. When Facebook decides to take a smattering of our data and make it "Public" how then do they decide to control that data after the fact? Since when did Facebook re-define the word "Public"? It's like your cell phone provider and ISP redefining "unlimited". This stuff has to stop. Getting back to Facebook, since when does merely browsing to a URL(I) enter you into some sort of binding contract with the publisher?
"what happens when Facebook gets sued for this data exposure?"
And there are two avenues that might happen - investors suing as to loss of potential revenue, and user(s) suing over privacy violations.
Bear in mind, FB have already had class actions against them.. i don't think setting aside another 10m+ (conservative) just to defend yet another nuisance lawsuit is something they're going to be up for.
So, as Pete Warden described, they bullied him into shutting it all down with threats that they could just as easily lose in court - based on the principle that it's cheaper for them to sue Pete and lose cred amongst us than it would be for them to face a class action from users who are led to believe that their privacy has been violated by the data set via an ambulance-chasing lawyer.
Facebook should have instead invited Pete in for a chat and a job, but instead they took the full frontal lawyer bulldog approach. Sometimes that happens. Hopefully next time it wont- and i've already poked my friends at FB to raise the issue internally if possible -- and so should you.
There's a difference.
In that case, your data is gathered from the search engine and has nothing to do with facebook any more. And I doubt the search engine will sue you for using their service.
I'm glad he posted about it though, that legal issue regarding robots.txt is good to know.
http://www.newscientist.com/article/dn18721-data-sifted-from...
I'd imagine they want to deter people from following in my footsteps, so publicizing their actions serves as a warning.
If anything, I'm encouraged to start a publicly-available database based on crawling Facebook. Don't associate your name with it, crawl from your Ipredator account, and make sure your data is spread far and wide. Then they can sue the Entire Internet like the music companies are having so much success with.
Information wants to be free.
(Also, what's the worst that could happen? You get sued, and Facebook gets a few thousand bucks from your savings account and your used mattress? OH NOES.)
By the time they get a court order to take the data down, if they even bother, people will have already distributed it, and it'll be another 09:f9:11:02-type fiasco.
I hope he releases the code he used, although I wouldn't really blame him if he doesn't.
I think the full quote is something like "Information wants to be free. Information also wants to be expensive". As I recall, it was about the dichotomy between the fact that you could copy bits for free, and that information was very valuable.
or am i wrong?
Now can't i publish my thesis because it contains information i got from copyrighted books?