People like traffic to their actual domain, because it often means better search engine ranking in the future. On top of that, some websites serve ads, which means that traffic is proportional to revenue.
Well, legal and whatever domain the "you viewed the ads but they werent tracked so we don't get money" problem falls into. Stupid artificial barriers which are there because they solve an economic problem, that wouldn't exist in a perfect world. Game theory, maybe.
Caching it for when a site is overloaded is not ripping it off, just like sending copies of works to the British Library for archival purposes is not ripping the original work off.
Also, the library analogy doesn't work in copyright law. Copyright protects the act of copying, not the act of transferring an already-authorized copy of a work to someone else.
Perhaps a simple GET to the URL ever 60s. If you receive 5 non-200 responses in a row, then the HN link points to the archive.is version. Same method in reverse for bringing it back.
How would one request per minute from a single server DoS the site, let alone DDoS it?
- not everyone upvotes
- there's probably a substantial number of users who are content to read articles but have never had a login (I did for two years)
- social media amplification.