It seems routine to see a bunch of browser User-Agents from the same IP
utcc.utoronto.ca
utcc.utoronto.ca
There's no "normal" UA cardinality count that you can use for filtering. However, there may be other traffic characteristics that are useful; the shape of the distribution, for instance.
This makes things like DoS protection... Interesting as we can regularly have thousands of users legitimately using our stuff from the same IP concurrently.
I often am repulsed that servers reads user agent. It is often not used for compatibility reasons, but just to detect if you are running official browser, or if you are 'privileged user'. This leads, obviously, that some programs needs to use disguise, which would not happen if servers were friendly toward friendly bots. Imho there is a problem with bad-bots, not all bots, that take away your bandwidth, processing power, etc. etc.
No one providing a free resource online has any obligation to treat your requests in a particular way. If you dont like the way a particular HTTP service is responding, then dont send it requests.
Or is mass crawling of static content completely fine and we should just be better at caching?
No. Bad actors generally don’t pay for CPU/GPU as they use compromised hosts or legit hosts with stolen payment info.
This same “proof of work” solution didn’t fix email spam 20 years ago, for the same reasons. It only adds costs to legitimate activities while being a minor inconvenience to bad actors.
Was this solution actually shown not to work 20 years because spammers are able to pay the cost? My sense is that Hashcash never saw widespread adoption, and we don't really know how it would have worked out.
I still use greylisting which keeps a remarkable number of spammers away, because they do not invest the resources to wait for n minutes for a retry. And I refuse email from clients without a proper reverse DNS entry, which more often than not appear to be hacked machines, or certain "organisations" in countries far away ...
I ran a website years ago that was targeted towards college students and would pretty consistently see many (hundreds) different UAs under the same IPv4 address due to NAT. Let alone proxies, VPNs, etc.
Yet every once and a while someone will suggest the idea of rate limiting based on IP so it's definitely not universal knowledge even among developers how common it is for multiple users to share the same IPv4.
Just a simple web search on issues parsing access logs would show this is not a new idea nor would it be a successful one.
I'm always surprised when a blog post like this makes it to the front page of HN. I suspect that their only experience with web hosting is this blog as there's nothing clever or interesting about their experiment.