But if these are popular apps / APIs, then the number of affected households is significant. Authorities / investigators will have to treat IPs as likely proxies and not the geolocation of the human initiating the request.
When you can identify the nature of the traffic (quickly in realtime, based on simple deterministic rules), you can protect the resources: you can rate/concurrency -limit the AI scrapers in the name of saving resources for the real humans, effectively putting the scrapers in a lower priority band (which is how it generally worked for search engine scrapers before!).
The problem is they're using resiproxies to disperse and whitewash their traffic, making it extremely difficult to tell their requests apart from the legitimate human requests. They're basically lying to us about the origin, and thus denying us the ability to put them in a lower priority band than humans.
They may scrape us at, say, 25K reqs/second, but it's coming from 50K random residential eyeball IPs at an average rate of only 0.5 reqs/second/IP, and then they're intentionally lying with the UA and headers and other fingerprint details as best they can to "blend in" with the humans so that we can't differentiate.
Let's do an analogy: Imagine if there was a neighborhood grocery store you and all your neighbors rely on for food. It's cheap because they keep their margins low, and more importantly the next store down the road is like 50 miles further away. That store 50 miles down the road also charges double the price. Now they've decided to play arbitrage: they load up 100 employees in the back of an air conditioned semi, clothe them to look like local shoppers, park it 3 blocks from your neighborhood store hidden inside a fenced property, and have them all go in and buy out all the inventory in the store over the course of a couple hours. The store just looks like it's having a great sales day at first. All these customers waiting in line, each getting just a few things at a time. But two hours later, the store shelves are empty, the semi is loaded up, and they're headed 50 miles back to double the price and sell it to someone else. You go in to buy some veggies to cook dinner and there's nothing to buy.
We've been playing this game with AI scrapers and resiproxies for way too long, and someone needs to hold them accountable for their fraud.
Who's using VPNs and Tor to blend in their automated scraping traffic with real human traffic?
Who's using multiple VPNs or Tor exit nodes to avoid rate limits?
No one, but I would have no problem with that being illegal too.
Tor less so but it doesn’t seem to be commonly used for this kind of abuse.
I expect AT&T and Comcast to offer a residential proxy service any day now.
Bear in mind the scrapers wouldn’t need to use these proxies were they not being blocked by the sites they are scraping. So it’s being used to evade blocks.
For some content the level of scraping is outweighing real users, driving up costs and pushing them towards more closed models.
Wikipedia for example make content available free, if you start hammering the site they will rate limit you to keep the lights on. If you need the data fast in bulk they have a paid program to get it without scraping. But some prefer to neither adhere to reasonable request limits nor pay for their use of the infra; instead they choose to pay these grifters to avoid the rate limits.