Archiving URLs (2019)
gwern.net
gwern.net
I am still finetuning it: there are a lot of domains which should be whitelisted because archiving doesn't work or isn't useful, and sometimes pages break when SingleFile'd (the content is there, but the CSS or JS is broken, and I muck around manually trying to fix them). The resource requirements are not that bad so far (21GB) but are probably a deal-breaker for anyone who wants the usual personal-static-site budget strategy of hosting on Amazon S3 for <$5/monthly, as even assuming way fewer links than gwern.net, cloud bandwidth is so expensive they'll quickly blow their budget. It also makes server logs messy as various bots try to fetch broken relative links, which were of resources which didn't get inlined by SingleFile (not sure what exactly triggers those).
However, so far it seems like a viable strategy.
It's probably something simple like "lots of websites encode addresses with absolute paths, so bots follow those; SingleFile ought to rewrite absolute paths to point to the original domain."
To me it's as if I lived on 15 Main St., and for that reason alone I had the right to block the publication of photo albums of past years years of 15 Main St., even if the dates are clearly indicated.
What they do is certainly tortuous infringement in UK - and probably most jurisdictions AFAICT (based on Berne Treaty, say), but it's so easy to remove your domains from IA that courts will be loathe to award much in damages.
AIUI they block access but don't delete the info; Fair Use/archival laws in USA probably make that lawful.
Second law of the Internet: you cannot keep track of where something is on the Internet.
https://github.com/pirate/ArchiveBox/wiki/Web-Archiving-Comm...
But alas, here we are.