Web scraping is legal, US appeals court reaffirms
techcrunch.com
techcrunch.com
This is an unpersuasive argument because it ignores all the computer users who are not "members". Whether or not "members" trust LinkedIn should have no bearning on whether other computer users who may or may not be "members" can retrieve others' public information.
Even more, this statement does not decry so-called scraping only "unauthorised" scraping. Who provides "authorisation". Surely not the LinkedIn members.
It is presumptuous if not ridiculous for "tech" companies to claim computer users "trust" them. Most of these companies recieve no feedback from the majority of their "members". Tech companies generally have no "customer service" for the members they target with data collection.
Further, there is an absence of meaningful choice. It is like saying people "trust" credit bureaus with their information. History shows these data collection intermediaries could not be trusted and that is why Americans have the Fair Credit Reporting Act.
They are spin doctors lying through their teeth, because literally everyone is. The culture of business writing/speaking has become so far disconnected from physical reality that it is amazing that anyone accepts these explanations at all, ever.
Like seriously, these guys have proferring casually miquetoast explanations for anything down to a science. If there is one thing you can guarantee, it's that the provided explanation simply isn't true.
It's propaganda and it would be great if it didn't work, but it works at least partially. Someone who doesn't know anything about the subject will hear it and repeat it, someone will hear the repetition and give it credibility because it doesn't come directly from a company. I have heard the most stupid opinions from people that go directly against their interests. You don't need fancy algorithms to manipulate society, just perseverance.
To be honest, a lot of people seem to want to kiss ass for the “prestigious” companies, too. You can see this phenomenon on all Apple threads, for example. They either don’t want their status symbols tainted, or they subconsciously believe that flattery will net them something.
The execs, lawyers, judges, public relations crews, and others who deal with the language day in and day it must become inured to it. Joe Public might never read an EULA, and might be shocked (if not already pessimistically resigned) to understand what companies actually do with his data, but an arbitrator who hears 6 complaints a day has been sitting in the boiling pot for a long time.
Finally, LinkedIns asserted private business interests protecting its members data and the investment made in developing its platform and enforcing its User Agreements prohibitions on automated scraping are relatively weak. ... Further, there is evidence that LinkedIn has itself developed a data analytics tool similar to HiQs products, undermining LinkedIns claim that it has its members privacy interests in mind."
> This is an unpersuasive argument because it ignores all the computer users who are not "members".
It's not just unpersuasive, it's disingenuous. LinkedIn wants to reap the benefits of having a public website while excluding competitors from their definition of "public".
If LinkedIn wants a membership-only website, they can privatize it like Facebook.
As a data scientist, I won't use LinkedIn after seeing them pursue this. They need to learn what public and private mean, and that it isn't the job of courts to punish people or businesses for accessing publicly available information. LinkedIn can set up its own boundaries as many other services already do.
Linkedin is being completely self serving. Taking whatever they can get, and feeding bullshit, sorry PR, to confuse the issue.
"We're disappointed, but this was a preliminary ruling and the case is far from over," a company spokesperson said. "We will continue to fight to protect our members' ability to control the information they make available on LinkedIn."1
1. http://www.theregister.com/2022/04/19/scraping_public_data_l...
Members cannot control the agreements that Microsoft/LinkedIn enters into with member information as the bargaining chip. There are no generally limits on how Microsoft/LinkedIn can use the information, either internally or externally.
I still get occasional emails despite having deleted my account years ago, after having a profile with them for a short time.
I've heard plenty of other dark pattern anecdotes about LinkedIn, and can assure them I would trust them as far as I could throw them.
LinkedIn presumably tells its users how they are using the data, at least if they follow the law. Shouldn't people be allowed to consider those terms to be acceptable without it meaning they lose all protection?
"But it's impossible to protect your data against all re-use", you'll say, "someone may remember it". That is not just obviously true, it is the only reason one might want the protection of the law.
Things that can easily be prevented by technical means do not need to be prohibited. That's why the common argument that "it's your own fault when you are raped at night, in a park" or that someone stole your car when you accidentally didn't lock it is so absurd: laws aren't so much about preventing you from being raped. In the absense of the law, you'd just never leave your (fortified) appartment. Laws are about allowing you to go outside, to not spend your money on steel-reinforced doors, to leave your convertible parked with the top down.
"Might is right" just leads to a pointless arms race: social networks waste money on protecting against scrapers. They'll hide everything behind logins or paywalls. Scrapers will waste money on overcoming those protections.
If the scrapers win, some people will decide not to do something they would have otherwise done. In other words: they have become a bit less free. The social network becomes less useful and might shut down. The scrapers find themselves with nothing to scrape. Congratulations, everone lost.
The ruling specifically says that scraping data that is publicly accessible (which I presume means without login) is OK.
LinkedIn serves pages containing their users' info, and users are made aware of this. These pages require no authentication, presumably because they're a marketing tool. People built scrapers to obtain that public info. LinkedIn said that was illegal, the court says it isn't.
I don't see who loses here, other than LinkedIn and other sites that want the benefits of listing information without the downsides.
Edit: Would it be illegal photograph the board, OCR it and provide a summary for half the price of others? There is such a thing and "wrong/stupid laws"
I think the real action will be the ruling of the district court.
[1] https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/1...
[1] https://storage.courtlistener.com/recap/gov.uscourts.cand.31...
Second Thought: I should really read the opinion before I make any comments.
Clicked through, skimmed around, looked for a link to the slip looking for the opinion link ... couldn't find it.
Thank you, sir.
For the family of sites I'm responsible for, bot traffic comprises a majority of traffic - that is, to a first approximation, the lion's share of our operational costs are from needing to scale to handle the huge amount of bot traffic. Even when it's not as big as a DOS, it doesn't seem right to me that I can't tell people they're not welcome to cause this additional system load.
Or even if there was some standardized way that we could provide a dumb API, just giving them raw data so we don't need to incur the additional expense of the processing for creature comforts on the page designed to make our users happier but the bots won't notice.
I'll skip the details, but a previous employer dealt with a large, then-new .mil website. Our customers would log into the site to check on the status of their invoices, and each page load would take approximately 1 minute. Seriously. It took about 10 minutes to log in and get to the list of invoices available to be checked, then another minute to look at one of them, then another minute to get out of it and back into the list, and so on.
My job was to write a scraper for that website. It ran all night to fetch data into our DB, and then our website could show the same information to our customers in a matter of milliseconds (or all at once if they wanted one big aggregate report). Our customers loved this. The .mil website's developer hated it, and blamed all sorts of their tech problems on us, although:
- While optimizing, I figured out how to skip lots of intermediate page loads and go directly to the invoices we wanted to see.
- We ran our scraper at night so that it wouldn't interfere with their site during the day.
- Because each of our customers had to check each one of their invoices every day if they wanted to get paid, and we were doing it more efficiently, our total load on their site was lower than the total load of our customers would be.
Their site kept crashing, and we were there scapegoat. It was great fun when they blamed us in a public meeting, and we responded that we'd actually disabled our crawler for the past week, so the problem was still on their end.
Eventually, they threatened to cut off all our access to the site. We helpfully pointed out that their brand new site wasn't ADA compliant, and we had vision-impaired customers who weren't able to use it. We offered to allow our customers to run the same reports from our website, for free, at no cost to the .mil agency, so that they wouldn't have to rebuild their website from the ground up. They saw it our way and begrudgingly allowed us to keep scraping.
In reality it was probably more like org sub group A wanted to leverage org sub group B’s data but they didn’t cooperate
I understand that bots have leverage and automation, but so does you to reach a larger audience. Should we continue to benefit from one side of the leverage, while complaining about the other side?
My complaint is more like, somebody wants to know the prices of all our products, and that we have roughly X products (where X is a very large number). They get X friends to all go into the store almost simultaneously, each writing down the price of the particular product they've been assigned to research. When they do this, there's scant space left in the store for even the browsing kind of customers to walk in. (of course I exaggerate a bit, but that's the idea)
That said if scraping is inevitable, it’s immensely wasteful effort to both the scraper and the content owner that’s often avoidable.
Also, I'm a big Jenson Button fan.
Would someone let me know if I’m just plain wrong in this assumption? I’ve run many types of sites and scrapers have never been anywhere close to the main source of traffic or even particularly noticeable compared to regular users.
Even considering a very commonly scraped site like LinkedIn or Craigslist - for any site of any magnitude like this public pages are going to be cached so additional scrapers are going to have negligible impact. And a rate limit is probably one line of config.
I’m not saying you are necessarily wrong, but I can’t imagine a scenario that you’re describing and would love to hear of one.
But the bots on my site -- at least the obvious ones that lead me to say they are a large source of traffic -- are all well-behaved, with good clear user-agents, and they respect robots.txt, so I could keep them out if I wanted.
I haven't wanted because, why? I have modified the robots.txt to keep the bots out of some mindless loops trying every combination of search criteria to access a combinatorial expansion of every possible search results page. That was doing neither of us any good, was exceeding the capacity of our papertrail plan (which is what brought it to our attention) -- and every actual data page is available in a sitemap that is available to them if they want it, they don't need to tree-search every possible search results page!
In some cases I've done extra work to change URL patterns so I could keep them out of such useless things with a robots.txt more easily, without banning them altogether. Because... why not? The more exposure the better, all our info is public. We like our pretty good organic Google SEO, and while I don't think anyone else is seriously competing with google, I don't want to privilege google and block them out either.
Bots would routinely try to scrape pricing for every combination of {property, arrival_date, departure_date, num_guests} in the next several years. The load to serve this would have been vastly higher than real customers, but our frontend was mostly pretty good at filtering them out.
We also served some legitimate partners that wanted basically the same thing via an API... and the load was in fact enormous. But at least then it was a real partner with some kind of business case that would ultimately benefit us, and we could make some attempt to be smart about what they asked for.
As described elsewhere, rate limiting doesn't work. The bots come from hundreds to thousands of separate IPs simultaneously, cooperating in a distributed fashion. Any one of them is within reasonable behavioral ranges.
Also, caching, even through a CDN doesn't help. As a B2B site, all our pricing is custom as negotiated with each customer. (What's ironic is that this means that the pricing data that the bots are scraping isn't even representative - it only shows what we offer walkup, non-contract customers.) And because the pricing is dynamic, it also means that the scraping to get these prices is one of the more computationally expensive activities they could do.
To be fair, there is some low-hanging fruit in blocking many of them. Like, it's easy to detect those that are flooding from a single address, or sending SQL injection attacks, or just plain coming from Russia. I assume those are just the script kiddies and stuff. The problem is that it still leaves a whole lot of bad actors once these are skimmed off the top.
[1] https://en.wikipedia.org/wiki/List_of_largest_Internet_compa...
Being on that list puts the company's revenue at over $1 billion USD. At a certain point it becomes cheaper and easier to fix the system to handle the load.
So your company is deliberately trying to frustrate the market, and doesn't like the result of third parties attempting to help market efficiency? It seems like this is the exact kind of scraping that we generally want more of! I'm sorry about your personal technical predicament, but it doesn't sound like your perspective is really coming from the moral high ground here.
No. First, we as a middleman resseller MUST provide custom prices, at least to a certain degree. Consider that it's typical for manufacturers to offer different prices to, e.g., schools. This is reflected by offering to us (the middleman) a lower cost, which we pass on to applicable customers. Further, the costs and prices vary similarly from one country to another. Less obviously, many manufacturers (e.g., Microsoft, Adobe, HP) offer licensing programs that entitle those enrolled to purchase their products at a lower cost. So if nothing else, the business terms of the manufacturers whose products we sell necessitates a certain degree of custom pricing.
Second, it seems strange to characterize as "frustrating the market" what we're doing when we cooperate with customers who want to structure their expenses in different ways - say, getting a better deal on expensive products that can be classified as "capital expenses" while allowing us to recover some of that revenue by charging them somewhat more for the products that they'd classify as operational expenses.
> it seems strange to characterize as "frustrating the market" what we're doing when we cooperate with customers who want to structure their expenses in different ways
I'm characterizing the overall dynamic of keeping market price discovery from working as effectively. How you may be helping customers in other ways is irrelevant.
They’re trying to pay less tax by convincing your company to put a different price on products they buy based on their tax strategy.
It sounds illegal.
What do you think about their other assertion that the search page is getting a gigantic number of hits that a/ cannot be cached and b/ cannot be rate limited because they're using a botnet?
The scale of the botnet sounds like an awfully determined and entrenched adversary, likely arising because this company has been frustrating the market for quite some time. A good faith API wouldn't make the bots change behavior tomorrow, but they certainly would if there were breaking page format changes containing a comment linking to the API.
The thing I still don’t understand is why (edit server not cdn) caching doesn’t work - you have to identify customers somehow, and provide everyone else a cached response at the server level. For that matter, rate limit non-customers also.
Search results obviously can't be cached, as it's completely ad hoc.
Product details can't be cached either, or more precisely, there are parts of each product page that can't be cached because
* different customers have different products in the catalog
* different products have different prices for a given product
* different products have customer-specific aliases
* there's a huge number of products (low millions) and many thousands of distinct catalogs (many customers have effectively identical catalogs, and we've already got logic that collapses those in the backend)
* prices are also based on costs from upstream suppliers, which are themselves changing dynamically.
Putting all this together, the number of times a given [product,customer] tuple will be requested in a reasonable cache TTL isn't very much greater than 1. The exception being for walk-up pricing for non-contract users, and we've been talking about how we might optimize that particular cases.
The low millions of products also makes some sense I suppose but it's hard to imagine why this doesn't simply take a login for the customer to see the products if they're unique to each customer.
On the other hand, I suspect the price this company is paying to mitigate scrapers is akin to a drop of water in the ocean, no? As a percent of the development budget it might seem high and therefore seem big to the developer, but I suspect the CEO of the company doesn't even know that scrapers are scraping the site. Maybe I'm wrong.
Thanks again for the multiple explanations in any case, it opened my eyes to a way scrapers could be problematic that I hadn't thought about.
I would think that artificially slowing down search results can discourage part of the bots. Humans don't care much it a starch finishes in 5 seconds and not 2 AFAIK.
Especially on backends where each request is a relatively cheap operations wise (especially when each request is a green thread like in Erlang/Elixir), I think you can score a win against the bots.
Have you attempted something like this?
I haven't had to defend against extensive bot scraping operations -- only against simpler ones -- but I've utilized such a practice in my admittedly much more limited experience, and was actually successful. Not that the bots gave up but their authors realized they can't accelerate the process of scraping data so they dialed down their instances, likely to save money from their own hosting bills. Win-win.
Apologies, I don't mention to lecture you, just sharing a small piece of experience. Granted that's very specific to the backend tech but what the heck, maybe you'll find the tidbit valuable.
It's been forever since I worked at Yahoo Travel, but bot traffic was significant then, I'd guess roughly 5-10% of the traffic was declared bots, but Yandex and Baidu weren't agressive crawlers yet, so I wouldn't be terribly surprised if a site with a large catalog that wasn't top 3 with humans would have a majority of traffic as bots. For the most part, we didn't have availability issues as a result of bot traffic, but every once in a while, a bot would really ramp up traffic and cause issues, and we would have to carefully design our list interfaces to avoid bots crawling through a lot of different views of the same list (while also trying to make sure they saw everything in the list). Humans may very well want to have all the narrowing options, but it's not really helpful to expose hotels near Las Vegas starting with the letter M that don't have pools to Google.
Pretty much Twitter and the majority of such websites.
It seems like it should be perfectly legal to detect and then hold the connection open for a long period of time without giving a useful response. Or even send highly compressed gzip responses designed to fill their drives.
Legal or not, I can’t see any good reason that we can’t make it painful.
We all benefit from open data. Polite scrapers are just fine and a natural part of the web ecosystem.
Google has been scraping the web all day every day for decades now.
However presumably all the other provisions of the CFAA still apply, so if your scraping damages the functioning of a internet service then you still would have committed the crime of "Damaging a protected computer by intentional access". Negligently damaging a protected computer is punishable by 1 year in prison on the first offence. Recklessly damaging a protected computer is punishable by 1-5 years on the first offense. And intentionally damaging a protected computer is punishable by 1-10 years for the first offense. These penalties can go up to 20 years for repeated offenses.
Have you tried making your UI more challenging to scrape and adding a simple API that requires free registration?
I work in E-Commerce and (needless to say) we scrape a lot of websites. Due to our growth and the increase in scrapers we require, I’ve been writing a proposal to a higher up to talk to our biggest competitors to all set up a public API that batches the data for a smaller amount of requests.
It would save everyone quite some traffic and effort.
First, as a B2B site, many of our users from a given customer (and with huge customers, that can be many) are coming through the same proxy server, effectively presenting to us as the same IP,
Second, the bots years back became much more sophisticated than a single, or even relatively finite, IP. Today they work AWS, Azure, GCP, and other cloud services. So the IPs that they're assigned today will be different tomorrow. Worse, the IPs that they're assigned today may well be used by a real customer tomorrow.
It obviously depends on how motivated the scrapers are (i.e. whether their headless browsers are actually headless, and/or doing everything they can to not appear headless, whether Google has caught on to their latest tricks etc. etc.) but it would at least be interesting to look at the score distribution and then see whether you can cut off or slow down the < 0.3 scoring requests (or redirect them to your API docs)
If someone dedicated themselves to it, there’s a lot more that these solutions could be doing to distinguish between humans and bots, but it requires true specialized talent and larger expenses.
Also, for a handful of the companies which make the most popular captcha solutions, I don’t think the incentives align properly to fully segregate human and bot traffic at this time.
I think we’re still very much still picking at the very lowest hanging fruit, both for anti-bot countermeasures and anti-anti-bot (counter-countermeasures).
Personally I believe this will finally accelerate once AI’s can play computer games via a camera, keyboard, and mouse. And when successors GPT-3 / PaLM can participate well in niche discussion forums like HackerNews or the Discord server for Rust.
Until then it’s mainly a cost filter or confidence modification. As long as enough bots are blocked so that the ones which remain are technically competent enough to not stress the servers, most companies don’t care. And as long as the businesses deploying reCAPTCHA are reasonably confident that most of the views they get are humans (even if that belief is false), Google doesn’t have a strong incentive to improve the system.
Reddit doesn’t seem to care much either. As long as the bots which participate are “good enough”, it drives engagement metrics and increases revenue.
just what kind of scraper you have is a concern.
does scraper just want a bunch of stock images;
or does scraper have FOMO on web trinkets;
or does scraper want to mirror/impersonate your site.
the last option is the most concerning because then;
scraper is mirroring bcz your site is cool and local UI/UX is wanted;
or is scraper phishing smishing or otherwise duping your users.
I feel that ruling should have the caveat that if a fair cost paid API version for getting publicly listed data then the scrapers must legally use that (say no more than 5% more than cost of CPU/bandwidth/etc of the scraping behaviour); ideally a rule too that at minimum there be a delay if they are republishing that data without your permission, so at least you as the platform/source/reason for the data being up-to-date aren't harmed too - which may then kill the source platform over time if regular visitors somehow start going to the competitor publishing the data.
Here's the end of LinkedIn's robots.txt:
User-agent: * Disallow: /
# Notice: If you would like to crawl LinkedIn, # please email whitelist-crawl@linkedin.com to apply # for white listing.
You can tell them. You just can't prosecute them if they don't obey.
Like when Aaron Swartz spent months hammering JSTOR causing it to become so slow it was almost unusuable, and despite knowing that he was causing widespread problems (including the eventual banning of MIT's entire IP range) actually worked to add additional laptops and improve his scraping speed...all the while going out of his way to subvert MIT's netops group trying to figure out where he was on the network.
JSTOR, by the way, is a non-profit that provides aggregate access to their cataloged archive of journals, for schools and libraries to access journals they would otherwise never be able to afford. In many cases, free access.
I'm surprised to see someone so cold and unfeeling about Aaron Swartz. Especially considering the massive injustice with regards to application of the law and sentencing.
> Federal prosecutors, led by Carmen Ortiz, later charged him with two counts of wire fraud and eleven violations of the Computer Fraud and Abuse Act, carrying a cumulative maximum penalty of $1 million in fines, 35 years in prison, asset forfeiture, restitution, and supervised release.
If not, you should, and the badly-behaved scrapers are actually a good wake-up call.
The only library that I know that is more or less undetectable is used by a just a few hundred people...
We do use a 3rd party service to help with this - but that on its own is imposing a 5- to 6-digit annual expense on our business.
and you're sweating a 5- to 6- digit annual expense?
> all our pricing is custom as negotiated with each customer.
> there's a huge number of products (low millions) and many thousands of distinct catalogs
Surely the business model where every customer has individually-negotiated pricing model costs a whole lot to implement, further, it gives each customer plenty of incentive to attempt to learn what other customers are paying for the same products. Given the tiny costs of fighting bots, in comparison, your complaints in these threads here seem pretty ridiculous.
those are only the low-effort/cheap ones, the more advanced scraping makes use of residential proxies (peoples' pwned home routers, or where they've installed shady VPN software on their PC that turns them into a proxy) to appear to come from legitimate residential last mile broadband netblocks belonging to comcast, verizon, etc.
google "residential proxies for sale" for the tip of an iceberg of a bunch of shady grey market shit.
If you're dropping 6 figs annually on this and it's still frustrating, I'd be interested in talking with you. I built an abuse prediction system out of this approach for a small company a few years back, it worked well and it'd be cool to revisit the problem.
IANAL, but I also wonder if, given that I'd be designing something specifically for competitors to query our prices in order to adjust their own prices, this would constitute some form of illegal collusion.
Scrapers have a limited range of IPs, so rate-limiting them and stalling (or dropping) request responses is one way to deal with the DoS scenario.
For my sites, I have placed the majority behind HTTP Basic Auth...
[0] https://brightdata.com/proxy-types/residential-proxies [1] https://oxylabs.io/products/residential-proxy-pool
> Bright Data has built a unique consumer IP model by which all involved parties are fairly compensated for their voluntary participation. App owners install a unique Software Development Kit (SDK) to their applications and receive monthly remuneration based on the number of users who opt-in. App users can voluntarily opt-in and are compensated through an ad-free user experience or enjoy an upgraded version of the app they are using for free. These consumers or ‘peers’ serve as the basis of our network and can opt-out at any time. This model has brought into existence an unrivaled, first of its kind, ethically sound, and compliant network of real consumers.
I don't know how they can say with a straight face that this is 'ethically sound'. They have, essentially, created a botnet, but apparently because it's "AdTech" and the user "opts-in" (read: they click on random buttons until they hit one that makes the banner/ad go away) it's suddenly not malware.
Tesonet and other similar services (e.g. Luminati) don't have that. As far as anyone -- including web services, the ISP, or law enforcement -- are concerned, their traffic is the subscriber's traffic.
I would be interested in a reference for this if you have one.
If we were to analogize this to a non-internet example: (1) A company throws a free concert/event and believes they will make money by alcohol sales. (2) A bunch of sober/non-drinking folks attend the concert but only drink water (3) Company blames the concert attendees for "taking advantage" of them when they really just had poor company policies and a bad business model.
Put things behind authentication and authorization. Add a paywall. Implement DDOS and detection and banning approaches for scrapers. Etc etc.
But don't make something public and then get mad at THE PUBLIC for using it. Behind that machine is a person, who happens to be a member of the public.
That’s what it feels like when someone is scraping your network to bootstrap a competitor.
In school I needed to scrape a few hundred thousand pages of a proteomics database website. For some reason you had to view each entry one at a time. There was IP throttling which banned you if you made requests too quickly. But slowing the script to 1 request per second would have taken days to scrape the site. So I paid <$5 for a list of 500 proxy servers and distributed it, completing the task in under half an hour.
I’m surprised your school was okay with it.
And, it's absolutely worth the cost - as a website owner, you get to impose costs on botting operations with minimal penalties for normal users and minimal environmental impact. Bots work because the costs of renting an AWS server and scraping websites (or sending spam, whatever) are extremely tiny - adding PoW challenges to everything that could be spammed suddenly massively changes the cost of running those spam operations, and would result in noticeably less spam if deployed widely.
In fact, the net "environmental impact" would be negative, as botters start to shut down operations due to greatly increased operational costs.
This really is akin to the question, “Should others be allowed to take my photo or try to talk to me in public?”
Of course the answer should be yes, the internet is the digital equivalent of a public space. If you make it accessible, anyone should be able to consume.
If you don’t want it scraped add auth!
An old bureaucratic company developed a system for a local library. It was very slow, so a user started scraping it and developed his own library system. A few weeks later, he was arrested.
The system was so shit that it became unusable by only hundreds of requests per hour. The local government filed a claim, and the officer arrested him for the "dangerous cyber attack."
[0] In Japanese: https://www.wikiwand.com/ja/%E5%B2%A1%E5%B4%8E%E5%B8%82%E7%A...
If we assume the very high end of "hundreds", 900, it's still 15 requests per minute. Your system is a disgrace if it can't handle that kind of traffic. This incident was in 2010 so even that is not an excuse.
Website admins don't owe you anything. If their website sucks, it sucks. That doesn't give you the right to take down their system.
A DDoS is a particular kind of DoS attack - one in which the attack is conducted using a distributed set of machines, usually helping to (a) hide the attack from monitoring tools and (b) make it hard to block using IPs. Obviously this was not a DDoS, but could have looked like a plain DoS.
The definition of a DoS attack depends on the system being attacked, there is no hard and fast limit to how much traffic a system is supposed to handle. For a badly written system like the one described here, the actions of the scraper were indistinguishable from the point of view of the library from the actions of a malicious actor. Hopefully, since the intent was not malicious in this case, thing may have been cleared up - though of course it's also possible that, unfortunately, they haven't.
By your own definition, it was not a DoS attack, because the subject was not seeking to deny others access. And without the intent to do harm, one may not genuinely identify the incident as an attack. Requiring intent is an important part of your definition, as it protects innocent users of a buggy system from being classified as attackers whenever any downtime incident occurs.
The ideal thing in this case may have been for the police to investigate, to conclude that the DoS was not done with malicious intent; but also to instruct the scraper that any further attempts to scrape the system in a this way will be considered malicious.
In this case, after volunteer engineers invested in their system, it turned out that they created a DB connection after each request, maintained it for 10 mins, and didn't reuse it or close it. So the number of DB connections quickly reached its limitation and couldn't handle the following requests. Similar problems happened with other library systems developed by the same company.
The prosecutors dropped the charge citing his scraping was reasonable, and he didn't have any malicious intent.
1. The page is not behind a login. If you have to create an account to access the content, then you have to agree to abide by the terms of service which may ban scraping.
2. What you scrape is still protected by copyright. If you scrape a clip from Star Wars from a web page, you still can't redistribute it without a license.
3. Your activity may not impede access to the site in anyway i.e. no DDOSing. If there is a robots.txt file, you are supposed to abide by it but the courts do not literally say that. Must rate meter your requests, scrape at night, so on. Gray area if the robots.txt bans every page and you end up having to ignore it.
[0]: https://www.eff.org/deeplinks/2018/01/ninth-circuit-doubles-...
However, the case of scraping I personally find more problematic is the use of personal data I provide to one side, then used by scrapers without my knowledge or permission. I truly wonder which way we are better off on that issue as a society. Independent of the current law, should anything that is accessible be essentially free-for-all or should there be limitations on what you are allowed to do. Cases highlighted in the article: Facial recognition by third parties on social media profiles, facebook scraping for personal data, search engines, journalists or archives. (Not all need to have the same answer to the question "do we want this") Besides that, the point I care slightly less about is the idea that allowing scaping with very leisure limits leads to even more closed up systems.
There is such a thing as scraping responsibility and irresponsibly. Both kinds happen.
First, one question is if the intent of the original owner of the data important? When I put data on linkedin (or facebook, my private website, hackenews or my employers website) I might have an opinion on who gets to do what with my data (see also GDPR discussions). Should I blame linkedin (or meta/myself/my employer) to do what I expected them to do, or should I blame those that do what I don't want them to do? Should I just be blamed directly because I even want to make a distinction between those? If I didn't want my data I could just not provide it (or participate in/surf the web at all if we extend the idea to more general data collection).
Secondly, it touches on the idea that linkedin should not make the data publicly available (i.e., without authentication), and we end up with a less open system. Is that better? Is it what we want? Maybe there are also other ways that I am not aware just now. (Competing purely on value added is probably futile for data aggregators.)
As for whether the data should be public, that's a decision we each have to make.
If by "information" you mean mere facts without creativity in selection or arrangement, those are generally not protectable by copyright in the United States, although possibly in some other countries. Copyright generally protects works of authorship, and nothing else. No creativity no copyright.
I used to regularly see books at the university book store that were wrapped in literal shrink wrap licenses for that reason, which of course makes browsing difficult. And then there was a major case on the subject ProCD v. Zeidenberg (7th Cir 1996), where such a license was found valid and enforceable.
But also anything you write which is copyrightable is copyright immediately. You can register the work w/ the copyright office for some extra perks but it's not strictly necessary.
So they were specifically interested in you, personally not the aggregate
Still, a very good day for scraping.
It's only after they had become truly huge that they started actually serving data directly (news snippets, song lyrics, factual answers, maps, reviews etc).
Of course. The amount of the original work is one of the key criteria on whether something is fair use or not. And there have been debates over whether Google crosses that line or not.
But no one except the most ardent anti-copyright maximalists would argue that so long as something is available on a public web page, anyone can just reproduce it verbatim and distribute it on their own site. Scrapers do a lot of this anyway but doing anything about it is a wack-a-mole game of publications often don't bother.
https://patentlyo.com/patent/2020/08/court-pacer-should.html
But yeah, the irony of the federal court system legalizing screen scraping, something PACER contractually prohibits.
Court records should be public record, not encumbered in terms of use/access/reproduction/copyright.
https://www.eff.org/deeplinks/2013/07/weevs-case-flawed-begi...
This is kinda like telling someone who is being harassed on social media that they’re consenting to it by having a public account. We should strive to make our digital personae safe from bad actors, not throw our hands up and say “if you put yourself out there, you have no recourse”.
Again, I am pro-scraping. I see the benefits. But I don’t think this is a situation in which we have to accept all or nothing.
2 options.
Their Linkedin ID's are base 12, and would redirect you if you simply wanted to enumerate them.
You could also upload your 'contacts', 200-300 at a time and it'd leak profile IDs (Twitter and Facebook mitigated this ~5 years ago). I still have a @pornhub or some such "contact" that I can't delete from testing this.
Make it easier to get the data through an API than having to scrape it.
[0] https://www.eff.org/deeplinks/2013/07/weevs-case-flawed-begi...
Does "publicly accessible" also apply to websites that are totally free, but require registration (which is open to anyone)?
And even though this is in the US, do cases like these set precedents for other western nations?
asking for a friend ...
What's important here though is to confirm that this also can't somehow be waivered by "Terms Of Service". The law needs to clearly say that web scraping is legal and that sites can't just circumvent the law by agreement with users.
Even better if it could also clearly state that it includes any and all digital content and is not limited to (say) hypertext. I.e. if I as a user can reach content in any way, then I can also scrape/save/archive/whatever using the tech I want. E.g. I can save my youtube videos, spotify songs and so on.
That's why I usually recommend using US-based web scraping tools and services (such as https://webscraping.ai).
In other countries, it still may be a grey area.
Contemplating if I should switch phone number now.
The flaw in that argument is that a law is never needed, by definition, to prevent things that can easily be avoided by technical means.
The laws against burglary protect your property and health but, more importantly, they allow us all to live in peace without investing vast amounts of resources into physical protection.
I actually built an open source project for scraping real estate websites several years ago: https://github.com/RealEstateWebTools/property_web_scraper
Might be time to go back and update it ;)
https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/1...
1. It's hard to distinguish scrapers from DoS attackers, which we do expect to be illegal.
2. Depending on the use of the data after scraping, and the nature of the data being scraped, you may expect copyright laws and things like the GDPR to prevent any possible use of the data, so the very collection of it may be suspect.
3. Even for public data, a site may have Terms of Service that don't permit scraping, and some may expect to be able to enforce these