Fair. Scrapers should be polite and do their utmost to consume the smallest possible amount of resources.
> without my consent by masking your user-agent
Your consent is not required. It's my user agent. I set it to whatever I want.
> for the purpose of stealing data I didn't authorize you to have
Data can't be "stolen", only copied.
You set up an HTTP server that literally sends people the data when they request it. Don't do that if you don't want people to have the data. Secrecy is the only possible defense.
By the same argument I could say: If I send you an exploit and you execute it, don’t complain that your setup fell for it. Just don’t download and run random data from the internet.
In reality there’s a consent and expectation beyond the pure technicals.
I don't. I go out of my way to filter everything. Scraping is but one of the tools I use to do it. I want just the data that I actually care about, not people's javascripted hot mess websites full of malware-vectoring ads, fingerprinting and tracking.
I don't let my computers talk to strangers either. My servers don't respond to just anyone, they only reply to me, and only after I've cryptographically authenticated. When others try to talk to them it's like they're not even there.
But people want their computers to talk to strangers, don't they? They want to serve pages and pages of ads to massive audiences. Unlike your exploitation example, nobody's actively invading their computers and exfiltrating data. Breaking into someone else's computers and dumping their private databases is one thing. We're just requesting the exact same data that they're more than happy to send out to literally anyone who shows up with a browser, through the exact same channels even. So I really have no sympathy.
Here's some stuff on Adversarial Interoperability, which is an incredibly good thing that can reverse enshittification if it becomes more widespread: https://www.eff.org/deeplinks/2019/10/adversarial-interopera...
Though, scrapers can certainly steal capacity through conversion for their own use. When they do so, they permanently deprive the site owner and other users the beneficial use of that capacity at that time.
I'm speaking morally, not legally. Though, in the US at least, there are parallels with free newspapers. You're allowed to take one for free. It's not legal to clear the whole rack.
There are additional methods I chose not to document such as limiting access to logged in accounts that require double-opting-in to acceptable use policies and terms of use, not that most scrapers would give a toss. That it too much whack-a-mole for me personally. That method requires progressively adding friction to account creation and that comes with some pros and cons.
[1] - https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Som...
It's like public photography, it's intrinsically legal, except when it's a Flock camera and then it's suddenly an invasion of privacy.
The moment you openly publish information on the Internet, you have already given consent. There are other solutions to bandwidth usage.
User-agent discrimination should be illegal. All it does is further the control that Big Tech has, and help authoritarian governments with their control too.
Couldn't agree more. Nonsense like remote attestation too. If this discrimination persists, it will lead to the destruction of what little computing freedom we still have.
"If you didn't want me to do this, you should had a fence/cameras/security guards. You shouldn't have dressed like that. You shouldn't have put your phone in that pocket."
Excusing trillion dollar corporations like low level criminals is embarrassing. Society shouldn't have to lock itself up because bad actors are spreading everywhere. The bad actors should just be removed.
If these scraping companies would respect some kind of X-Scraping-Permitted header, this whole process would be a lot better, but that's not going to happen in an industry that has already normalized using botnets.