Captchas in general are a bit of an illusion really. The best Captcha you can have is a home-made one as the attacker has to go through actual manual effort rather than enabling --solve-recaptcha flag on the bot script.
Captchas in general are a bit of an illusion really. The best Captcha you can have is a home-made one as the attacker has to go through actual manual effort rather than enabling --solve-recaptcha flag on the bot script.
All you "protect" yourself from is casual script users and script kiddies which really can be solved by IP rate limit. If someone has access to thousands of IPs they can probably afford to drop $5 to solve the captchas too, right?
If there wasn’t, people wouldn’t use captchas
$5 per request is not a negligible amount of money. In practice it doesn't cost anywhere near that amount to call a MechanicalTurk API which will solve ReCaptcha for you. But it's still significant for any nontrivial number of requests, such as in the use case of scraping.
You should adjust your priors here. You're focused on the narrow case where a win condition is achieved by spending n dollars to solve a single instance of ReCaptcha. People who use ReCaptcha are (in my professional experience) overwhelmingly more focused on requiring ReCaptcha to be solved for every individual request of a given type.
I have been in the position you speak of, where I had a revolving set of IP addresses, requesting servers and user agents, and $5 per request would have immediately shut my operation down. As it was, the actual ~$0.15 per request to solve ReCaptcha was sufficiently significant that I couldn't curate enough data for what I needed, despite having all the other resources you mention.
It's dead easy to get around captchas unless you're just a casual scripter that wants to `wget` that one article - then fuck that guy, right?