What's the endgame here? AI can already solve captchas, so the arms race for bot protection is pretty much lost.
What's the endgame here? AI can already solve captchas, so the arms race for bot protection is pretty much lost.
Can defenses be good enough it's better to not even try to fight? It's a far harder question than wondering if a random bot can make a dozen requests pretending to be human
Make it easier to get the data, put less roadblocks in the way for legitimate access, and you'll find fewer scrapers. Even if you make scraping _very_ hard, people will still prefer scraping if legitimate use is even more cumbersome than scraping, or you refuse to even offer a legitimate option.
Admittedly, we are talking here because some people are scraping OSM when they could get the entire dataset for free... but I'm hoping these people are outliers, and most consume the non-profit org's data in the way they ask.
For example, I have a project that crawls the SCP Wiki (following best practices, ratelimiting, etc). If they were to restrict the API that I use it would break the website for people, so if they do want to limit the access they have no choice but to instead put it behind some set of credentials that they could trace back to a user and eliminate the public site itself. For a lot of sites that's just not reasonable.
When I compare that to our current internet the first thought is "but that won't scale to the whole planet". But the thing is, it doesn't need to. All of the problems I need computers to solve are local problems anyway.
For everyone else, this transition will not be a big deal (although your friends may ask you to occasionally spend a few cycles maintaining your part of a web of trust, because your bad decisions might affect them more than they currently do).
Websites previously would have their own in-house API to freely deliver content to anyone who requests it.
Now, a website should be a simple interface for a user that communicates with an external API and display it. It's the user's responsibility to have access to the API.
Any information worth taking should be locked away by Authentication - which has become stupid simple using oAuth w/ major providers.
So these people trying to extract content by paying someone or using a paid service should rather use the API which packages it for them and is fairly priced.
Lastly, robots.txt should be enforced by law. There is no difference from stealing something from a store, and stealing content from a website.
AI (and greed) has killed the open freedoms of the Internet.
Also maybe the recent rise in captcha difficulty is not companies making them harder to prevent bots but rather bots twisting the right answer. As I know it captcha works based on other users' answers so if a huge portion of these other users are bots they can fool the alghorithm into thinking their wrong answer is the right answer.
This is a similar situation.
(Not sure if created by the admins or a 3rd party, but done once for many is better than overlapping individual efforts).
Our S3 bucket is thankfully supported by the AWS Open Data Sponsorship Program.
We've had good success with
- Cloudflare Turnstile
- Rate Limiting (be careful here, as some of these scrapers use large numbers of IP addresses and User Agents)
Require login, then verify the user account is associated with an email address at least 10 yrs old. Pretty much eliminates bots. Eliminates a few real users too, but not many.
this is not a solution if you want a public internet (and sites that don't care about the public internet already don't have a problem)
Someone still needs to pay for that traffic. If it gets too much for cloud flare or whoever, you’re gonna get the bill.
The generation process of taking the raw text and assembling the page around it is typically rather expensive for most CMS systems. Sure - it isn't theoretically expensive, but unless you want to engineer a CMS from scratch most people just pick one off the shelf and then end up having to pay the CPU time overhead of wordpress etc.
At best any email I have is 4 or 5 years old.
That's before even getting into how you'd possibly verify email adress age, especially without preventing self-hosting.
They generally look at data leaks and partner with big companies to see when that email address first signed up to any online service.