If you try and run a site that has content that LLMs want or expensive calls that require a lot of compute and can exhaust resources if they are over used the attack is relentless. It can be a full time job trying to stop people who are dedicated to scrapping the shit out of your site.
Even CF doesnt even really stop it any more. The agent run browsers seem to bypass it with relative ease.
This is wrong. Git does store full copies.
Prebuild statically the most common commits (last XX) and heavily rate limit deeper ones
2. 1M independent IPs hitting random commits from across a 25 year history is not, in fact, "easy to solve". It is addressable, but not easy ...
3. why should I have to do anything at all to deal with these scrapers? why is the onus not on them to do the right thing?
Is it pretty? No, but it also is a pretty niche thing overall (git repo storage).
The main advantage of Turnstile is that is benefits from CFs ubiquity to help judge legitimate vs illegitimate requests.
I would love to know what other options are available in this space aside from Turnstile, Recaptcha and HCaptcha.