Vast majority of websites today can and should be static, which makes even the aggressive llm scrapping non-issue.
This is wrong. Git does store full copies.
Prebuild statically the most common commits (last XX) and heavily rate limit deeper ones
2. 1M independent IPs hitting random commits from across a 25 year history is not, in fact, "easy to solve". It is addressable, but not easy ...
3. why should I have to do anything at all to deal with these scrapers? why is the onus not on them to do the right thing?
Is it pretty? No, but it also is a pretty niche thing overall (git repo storage).