IMHO, LinkedIn doesn't have a right to stop scraping after the fact, but they have the right to take technical steps to stop scrapers from accessing their site.
IMHO, LinkedIn doesn't have a right to stop scraping after the fact, but they have the right to take technical steps to stop scrapers from accessing their site.
That being said, I hate LinkedIn as a company and I fully support anyone trying to mess with them. They are not a social network, they are a sleazy website that convinces people to willingly provide personal information which they then turn around and sell at ridiculously high prices. Even if you are legitimately using LinkedIn as an end-user, it’s easy to get blocked for using it too much and being forced to pay just to interact with people on the site.
So it's fine to mess with them, even illegally, just because you don't like them?
* convinces people to willingly provide personal information
Convincing is not forcing, and in fact, you say "willingly" yourself. Any business convinces you to willingly give them money or other value.
* sell at ridiculously high prices
To respect to what? Prices are determined by the market. We are not talking about a life-saving medicine or health care, on which there could be some debate.
We are talking about a company selling information, which has value and which it acquires through an infrastructure that takes a lot of money to run. Value that customers are willing to pay for.
* Even if you are legitimately using LinkedIn as an end-user, it’s easy to get blocked
For what definition of legitimately? Yours? Since it's their business, they can define what is a legitimate free use and what can be a paid one.
* They are not a social network
So based on what I pointed out, they are not a social network only because they are not free? Do all social network have to be free and give your data for advertisers to be social networks?
If it's not illegal, then scummy behaviour against a scummy company doesn't exactly set my moral compass off. You reap what you sow.
The issue I see with this is in 2 part.
First, any issue that comes down to "moral compass" is inherently dangerous. We can find many examples of the simple concept that what to one person is Good is to another Evil. In this case, I think Linkedin shareholders would not appreciate calls to mess with the site, or with people having trouble jobhunting because the site is going down repeatedly due to DDOS or whatever.
The second is that these kinds of calls to action (linkedin sucks, fuck with it) smell like vigilantism to me, and while Batman is my favorite hero (I really only lift because I kinda sorta wanna be batman), vigilantism doesn't contribute to a stable society. Rule of law works better than the chaos of multiple agents enforcing their own moral code as law.
EDIT: I'm happy to be downvoted if I'm saying something stupid, but while doing so I would very much appreciate a quick comment as to why I'm wrong so I can improve my knowledge.
Most people seem to think this law (if it is held up in court) is wrong and should be changed.
You might say, oh well, we have a democratic right to change or influence our laws. But a princeton study has found no correlation between public preferences of the majority of the population and enacted policy: https://scholar.princeton.edu/sites/default/files/mgilens/fi...
From memory, the only thing that leads populations to revolt against their government is high enough food prices. Outside of that, revolts almost never happen.
What I'm trying to say is A) unjust, monopolistic, or excessive laws are probably more normal than the opposite because B) the idea that democracy means people have some weigh in lawmaking might be a myth and C) most people don't do anything about it because they only act when the very basics of their livelihood are threatened.
Your view, that some unjust or excessive laws are preferable to total chaos, seems to carry the assumption that laws are naturally benign and/or made to serve some purpose for society, therefore we should not challenge them without good reasons to do so. If the opposite is true and most laws or a high enough number of them are not just, then the fact that the vast majority of people disagrees with them is only natural.
This is a fairly long-winded way of saying that most people would say you're being downvoted because there are plenty of terrible laws that we should not acquiesce to silently.
Very good. That way the whole world will be blind and toothless. --Tevye, Fiddler on the Roof
If you allow yourself to act as badly as others act, you don't have much of a moral compass.
The world can only improve when we hold ourselves to a higher standard than we see, as it's too easy to rationalize our own behavior, and harshly judge others' actions.
The defense against a wolf is (generally) not eating the wolf.
The defense against invasion of privacy is not more privacy invasion.
The previous commenter never mentioned messing with them illegally and disagrees with the analysis that this scraping is illegal. He said he hopes the courts do not rule this is illegal. Pretty disingenuous to start a reply like that...
Many of your other points are not really giving the previous commenter any charity at all.
For example,
>For what definition of legitimately? Yours? Since it's their business, they can define what is a legitimate free use and what can be a paid one.
He was literally trying to provide an example of where their scraping protections can be appear to be overzealous to casual users, not arguing whether those users are legitimate or not.
What is the premise of your argument? To me, it seems you are simply trying to defend LinkedIn's business practices and legal pursuits, rather than discussing anything about the legality of scraping or the specifics of LinkedIn's anti-scraping implementations.
This probably happened a lot more a few years ago. Is perhaps 2FA making this harder these days?
I wonder if Google, Yahoo, MS etc have done anything like watch for requests from LinkedIn with correct credentials, block them, and reset the user's password and give them a warning that they just gave their account password to a third party and this is a Very Bad Idea.
I think that they used to have some FAQ entry explaining why worrying about this is silly and nothing bad could happen, but I can't find it any more (probably because it's nonsense). However, just because they should be shamed for this whenever possible, here's a Slate article on their overall security: http://www.slate.com/articles/technology/safety_net/2015/02/... .
Agree with you that giving your account password to any third party is madness, even more so actually soliciting it.
Literally yesterday I got, for the first time for that user, one of the spams sent "on behalf of" a user who clearly hadn't given out my e-mail address, so I guess they still are up to these no-good deeds.
What a shitty thing to do.
As far as I can tell, no one has made that argument, so I'm not sure why you feel the need to rebut it.
I think it all pretty much boils down to this quote from the article:
> LinkedIn's position disturbs Orin Kerr, a legal scholar at George Washington University. "You can't publish to the world and then say 'no, you can't look at it,'" Kerr told Ars.
The title is "It’s illegal to scrape our website without permission". So that argument is implied in the headline, at least.
As for Orin Kerr, I'm sure he'd agree that there are private parts of the internet (my payment information being an obvious example). Just because something is deployed to the internet doesn't mean it is "published to the world" as he claims.
If a company paid a million people to call into an information service hotline and each request one fact from it—and then the company recorded and compiled the answers into their own database to start their own information service—is that illegal?
I sure hope not. From there it doesn't seem to far off to claim that "He was really only ever showing up for work to learn some skills and how he's using those skills to run his own business!" sort of lawsuits. If you compile difficult to find but freely available information into a more easy to digest format I see that as virtually always a net positive.
He's not just bloviating; before you disagree with him, reviewing his arguments is worthwhile. (I do disagree with him in part, and agree with his reasoning but don't like the outcomes in part, but in any case, he's a pretty accomplished lawyer, and I'm not any kind of lawyer, so there's that.)
Everyone that scrapes LinkedIn (or anywhere else) either knows that they are doing it against LinkedIn's wishes or doesn't care.
I think you just defined a social network, unfortunately.
Every service has "terms of service" which are the conditions that you are allowed to access the service and what you may do with the service once being granted access. For example, if you start pouring toxic waste into your sewer, you will find that the city will both disconnect you from the service and they will fine you for violating the terms of service you nominally agreed to when being hooked up.
In LinkedIn's case, they allow you to access their service, with HTTP, to render a page in a browser for viewing of that page. Full stop. Any other use of the data you acquire over HTTP, or any other method of acquiring said data over HTTP is disallowed by the terms of service.
Not only does LinkedIn have a legal right to stop scraping after the fact, they have literally centuries of common law in support of that position.
With HTTP and LinkedIn, there is no such step. There's no pre-connection agreement. LinkedIn could present such an agreement on first connection, but they do not.
If you're talking about making anonymous requests to their service, they only allow a few of those before they stop showing you profiles. If you circumvent that protection, it's a bit more like hooking a cable up to a power line (illegal) or dumping your commercial waste in the sewer (illegal).
LinkedIn has two things that they do which protect them; First, they specify they disallow access in their robots.txt file. While not a binding agreement per se it is the default mechanism that is accepted by the community for apriori identifying whether or not automated access is possible. Second, when they detect an access pattern that violates their terms of service they actively block the access proactively notify the source of the violation.
The sad truth is that web scraping has been around since the very beginnings of the Web back in 1993 and this question has been litigated in every way that you might choose to argue it, the body of case law is enough to fill at least two volumes in the reference section of the library.
There is no legal or ethical basis for scraping the web without permission. And if it isn't explicitly allowed by a site the presumption is that it is disallowed (no 'open door' exception).
As for their ability to control what you do with the information: there might be a limited license on the data granted from users to LinkedIn that is not transferrable, so maybe you couldn't build a service that redistributed that information, but I don't see why obtaining and holding it would be illegal.
As for the analogies to power and telephone and such, those are built on property owned by a local government and there are usually other extra laws related to them: it isn't due to some common law position that you can't mess with their stuff. Here, I am not a lawyer, but I am a government official with a particular interest in sewage; here is a link to the sewer use ordinances form our local sanitation district: pay particular attention to 2.03.
http://goletawest.org/wp-content/uploads/2012/04/Ordinance-N...
Every single one of them concluded that based on how the law was written and how the web worked, there is no legal way to scrape a web site without its explicit permission to do so.
That won't stop people from trying of course and it was a source of constant entertainment in the ops team at Blekko at how people tried to sneak around at scraping (it can get very creative) but; it isn't legal, you can and will get banned from all access for it, and if you use the results in another product or offering you will be found liable for damages.
Google scrapes several of my sites and I've never given Google explicit permission to do so.
The implicit contract is that you let them scrape because you want to show up in their search results which will send you traffic. If you don't care about Google traffic then set /deny in your robots.txt and get back the bandwidth you were giving them.
Only for definitions of explicit I must be unfamiliar with.
If the presence of a robots.txt makes one's intent for a given resource explicit one way or the other, the lack of one (and the lack of some communication in some other channel) must mean there is no explicit permission.
> it is very clear to me that a search engine is
> operating on the legal equivalent of thin ice,
We may be saying similar things but from a metaphor I think of search engines operating on 'thick' ice. It has been litigated so much that there is a bevy of case law to refer to at all levels. Eric Goldman's blog used to have a pretty good list of the number of suits of various kind and the searchengine blog covered many of them as well.For a search engine it is super clear, robots.txt is all. If you say yes explicitly, great. If you say no explicitly, that has to be honored. If you say nothing, then its up to the search engine to decide which way to interpret it, but if the site owner complains because you picked wrong you have to honor their wishes (which may include destroying any cached data as well).
PadMapper, Perfect10, and the newspapers generated a ton of cases based on 'scraping a web site and using the data.' There are also about a dozen comparative shopping sites that have been dinged for the exact same issues. (look vs Amazon or vs Walmart).
Whether CFAA, DMCA, Torte law (contracts), or something else applies is constantly being discussed :-). I'm just the messenger here. I haven't found a single case that has held that the point of view of the scraper of someone else's web site should prevail. The argument that it should be allowed 'to help new businesses get off the ground' is like saying Apple should pay out some of its cash hoard as grants to startups trying to break into some business. I have yet to read anything that was sympathetic to that point of view.
But if you say, 'I am NOT a bot', like spoofing a browser's user agent string, but you are a bot, then you are requesting access under a pretense, in order to circumvent their terms of service. Kinda feels morally wrong, and illegal.
The point is, the scraper would have to hide their intentions and identity, which removes any claim they are being 'honest' in their intentions and not trying to circumvent the provider of the services efforts to prevent scraping.
Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Snackmaster Pro/666.0.666
What do you do?
I also tell my browser to lie about what it is sometimes, due to sites that are malfunctioning, but whose owners choose to document the errors instead of fixing them with "Use Chrome" (or IE, or whatever) checks.
Is that 'kinda' illegal or morally wrong (two very different things)?
If so, that seems like a belief that all sorts of browser defaults are 'kinda' wrong and/or illegal to change. Javascript? Lying about installed fonts/screen dimensions/whatever? Refusing to keep nonsession cookies between sessions? That slope would seem to get pretty slippery...
If by client you mean a robot, then you are pretending to be a browser and you are accessing the service without permission.
Let me ask you a question, say your client was hitting my service with that user agent, 100 times a second, crawling through urls sequentionaly. Lets say I added it to my robots.txt deny list and starting blocking that user agent. Would you change the user agent and continue?
If someone creates a site that says, 'Access to this site is for 640x480 browsers only, any other use is forbidden'. Then I think its pretty clear that its a stupid site but also that faking your screen resolution is accessing a site without consent. There is no slope, someone (Linkedin) putting explicit terms on their website is pretty clear.
What if I send a null UA? Or use it as an opportunity to share my favorite quote?
What if the behavior of my software doesn't attack like a robot, does keep the request volume reasonable (use whatever you think is reasonable here) but also doesn't do what you might expect a human clicking around to do?
- Reminds me of CraigsList vs PadMapper[1]. In that scenario I side with CL -- it was right to block PM. PM or others should not be allowed to build a new UI on top of CL because CL was the one that put in years of effort of nurturing its listings, its network, building brand equity and taking associated risks and costs.
- As others have highlighted, the data is publicly accessibly and there is no agreement the scraper/crawler is bound by. The agreement is between the LinkedIn user and LinkedIn. The scraper is connected to the Internet pipe crawling the Internet freely as it wants. It's not reproducing the data anywhere so copyright should not be an issue.
- What if a scraper didn't scrape LinkedIn but just the Google or Archive.org cached versions and read those instead? It would not be pressuring LinkedIn server resources in this case.
- What if all of my employees allow me to scrape their LinkedIn data? Can I scrape all of their info? Can LinkedIn stop me from doing that (In the case of Facebook vs Power Ventures, the answer is that LinkedIn would be able to prevent this behaviour).
- Who owns the data? Medium.com doesn't own the posts. LinkedIn doesn't own the CVs.
[1]: https://news.ycombinator.com/item?id=4286325I'm just saying the legal system doesn't see it that way, they have said so in many cases, and so far everyone who has used your argument or variations of it in court has failed to prevail.
# Notice: The use of robots or other automated means to access LinkedIn without
# the express permission of LinkedIn is strictly prohibited.
[1] https://www.linkedin.com/robots.txt"You agree that you will not ... develop, support or use software, devices, scripts, robots, or any other means or processes (including crawlers, browser plugins and add-ons, or any other technology or manual work) to scrape the Services."
https://www.linkedin.com/legal/user-agreement
That said, this should be a breach of contract issue. It's an overreach to invoke federal fraud law.
Most news sites publish to the world. But scraping a news site's content and monetizing it yourself is not ok. Legally, it violates intellectual property law. But laws aside, I assume most people would agree, if someone spent the time researching and writing an article, they should have the right to monetize it and nobody else.
In this case, IP law may not apply, but the concept is the same. I don't love LinkedIn myself. But they spent the time building a platform for collecting that info. I don't see why it should be OK for other people to scrape and monetize it.
EG, if it returns an image - it doesn't imply I can use the image anywhere I want.