ArchiveBox: Open-source self-hosted web archiving
archivebox.io
archivebox.io
https://github.com/ArchiveBox/ArchiveBox/wiki/Configuration#...
Not since 2017. https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
They now have a clunky manual process to exclude your site. https://help.archive.org/hc/en-us/articles/360004651732-Usin... ("How can I exclude...")
They don't spoof user agents, but blocking them actively doesn't remove their history.
The short version is that defaults in software are really important (90% of users wont change them), and I don't trust myself to code ArchiveBox 100% correctly so as to never lose data, or the majority of people to store their archives correctly so as to never lose data on their own. Archive.org is the redundant failsafe. Another good reason is that Archive.org is not the only way that your archive content can be leaked, the security model means that archived pages can read each other's content, so I want to make it abundantly clear to users that by default it's designed to only archive content thats already public (in which case it's already fair game for Archive.org).
I've settled on leaving it on as the default, but I do mention 3 times in the README how to disable it, most notably in the CAVEATS section which explains both the security model drawbacks and how to prevent your content from being leaked to Archive.org or other 3rd party APIs.
Nevertheless, no system is perfect, and even with Django helping guard database integrity and multiple redundant index files, it's possible I'll make a mistake someday that leads to data loss on upgrade. I don't want that situation to be the next (mini) library of Alexandria, and saving copies to Archive.org helps serve as a last-resort backup.
I personally have no issue with the defaults, but if you've agonized over the defaults, perhaps you should consider clearly documenting it in the main project README instead of leaving it for people to find in the config documentation.
FIY You just have to set the environment variable SUBMIT_ARCHIVE_DOT_ORG=False
Not uploading people's stuff to permanent, public archives seems like a good rule of thumb.
If it isn't already public (i.e., reachable by archive.org), it won't be afterwards?
I really wanna read this but can't find the thread, do you happen to have a link?
Just a heads-up. Found that a while ago and much prefer it over wallabag.
Any time I need to do anything, I will full clone the base; with a decent SSD it takes maybe 10 seconds for the full clone and I have a full OS.
For the very complex sites that really rely on a ton of interactive JS or dynamic requests to APIs to render their content, check out https://ArchiveWeb.page + https://ReplayWeb.page by https://webrecorder.io.
Could be very useful now.
./bin/export_browser_history.sh --safari
https://github.com/ArchiveBox/ArchiveBox/blob/dev/bin/export...https://stackoverflow.com/questions/28628385/sqlite-safari-h...
https://support.apple.com/guide/safari/keep-a-reading-list-s...