What are the legal issues around web scraping?
evan.law
evan.law
2. It is 100% legal to mask your IP address when accessing websites.
3. Always mask your IP address and you will never hear from the website owner.
(I'm not condoning anything illegal, and I have public-facing web-servers so I know all about the pain from aggressive bots and scraping. I'm just pointing out the facts around this instead of some vague lawyer's blog post with a "who knows, contact us" at the end of the article.)
From the website owner's perspective:
1. I get nailed with bots and scraping all day. If you do something outstandingly stupid and become the signal within the noise: I'm just going to harden my bot detection and mitigation instead of calling up my lawyer.
2. Almost all websites published today are running blind. They have no idea you are a bot, and they are more likely to get excited in the bump in traffic because "of course it's from the money we poured on that marketing effort".
3. If a website has a bot problem, they will recognize your scripts and give you guidelines on how not to be a jerk when scraping their content. This is usually backed up by IP based throttle limits.
4. I'm more afraid of this lawyer's hourly price, not the possibility of damage caused by web-scraping against my servers.
Random open web proxies are harder to defend against.
You can also spin-up your own server from a cloud provider and proxy your web traffic through that.
This is what was used against weev[1] when he crawled an AT&T website and dumped subscriber e-mail addresses, to sentence him to 41 months in prison. Don't get me wrong, weev is a terrible person that society should be protected from, but the severity of this sentence felt completely out of place for both the supposed "crime" and the banality of what he did.
If you could reasonably infer that you're not supposed to have access, try to obtain it and then go on to share that data with the media (which only acknowledges you know you're not supposed to have it) then I don't see a huge problem with prosecuting it.
Usually by scraping I understand accessing something that's already being accessed by clients and you're merely automating what's already happening.
Now if someone inadvertently obtained this kind of data without trying (say part of a bigger scrape) and didn't use or distribute it then obviously I don't want to see that prosecuted. And I doubt it would be.
I do, however, also recall weev saying how he spent most of his time in prison throwing nazi salutes and singing songs about white pride.
And if the author is going to bring up ToS and CFAA violation, it's negligent to not mention that recent high-profile precedents like LinkedIn vs HiQ exonerated the scraper of ToS and CFAA liabilities. Other precedents may say different for specific situations (pages behind a login?), but bringing up these legal obstacles with no indication that the best precedent we have makes them non-issues for scraping public information feels like the author just wants to scare the reader into thinking they need to call for a consultation.
delusional.
rhetorical language device intending to pretend you can make something public then dictate what public means to you.
If you're Google.
It's brilliant that the world's largest operator of web-scraping bots also develops the most-used bot-protection plugin (for reasons).