ArchiveBox: Open-source self-hosted web archiving
github.com
github.com
I encourage people to also check out the list of ArchiveBox alternatives we maintain if ArchiveBox doesn't quite fit your needs.
https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-...
Let me know if you find a good option later on and I'll add it to the list.
ArchiveBox also saves URLs it ingests to Archive.org by default for this reason!
Setting it up on K8s with sonic [1] as the search backend and importing a few hundred URLs only took ~an hour or so, and the cached pages look great for the most part.
So I wrote my own: https://github.com/tardisx/linkwallet
Emphasis on tiny system requirements and dependancies (single binary, no service dependencies). As a consequence the text indexing is very basic (basic HTML scrape). But it's working for me :-)
It's annoying because the site is autogenerated from the README markdown and it's tricky to add custom CSS without increasing build process complexity a bunch. PRs welcome!
For this purpose, I found the SingleFile browser extension to be the best fit. It's a browser extension, so paywall cookies are already present, and I just manually archive the previous week's content, after the discussion phase has concluded. It creates a single self-contained file with all images and comments, etc., but all non-page-local links still resolve externally (which is as-desired, for my use case). It can be configured to auto-generate a convenient filename, and to use self-extracting compression.
I preferred this to an automated process based on, e.g., RSS, because I can ensure the archive occurs after all the useful course comments back-and-forth has concluded, and it's trivial to set up and use.
ArchiveBox actually uses SingleFile internally as one our methods to save every page (among others), and we try to send a portion of our donations periodically to @gildas-lormeau to support his awesome work on it!
My primary concern about archivebox (and the WARC stuff) is the TB of existing archival stuff i already have.
I've been working on getting it deployed to fly.io with LSVD so it can scale to zero while storing everything on an S3-backed volume as described here[0].
My biggest disappointment so far is that it seems like a fairly large lift to make ublock origin work because extensions don't work in headless chrome (?). It seems like using pihole is current best method to block ads [1].
[0] https://community.fly.io/t/bottomless-s3-backed-volumes/1564... [1] https://github.com/ArchiveBox/ArchiveBox/issues/211
`--disable-extensions-except=/path/to/your/extension/` `--load-extension=/path/to/your/extension/`
We do add `ArchiveBox/v0.x.x` to the user agent for all requests by default + push URLs to Archive.org. So in theory someone at Archive.org could look in their server logs for that string and get a pretty good idea of the daily activity (at least for users with default settings). I've asked them a few times in person to run that search but never gotten a follow-up. They're probably very busy and it's just for curiosity, but it would be nice to know someday!
The only other metrics we have as of 2024/01:
- ~5m Docker Hub image pulls
- ~17k Github Stars
- ~1k issues and PRs, 100+ contributors
- ~1k browser extension users
Follow here for progress: https://github.com/ArchiveBox/ArchiveBox/issues/50
You can have multiple archives, and even use a mode where you only archive pages you bookmark rather than everything.
https://3xn.nl/projects/category/unraid/
First time user, but its one of those things I did not know I wanted.
note you no longer need to create a user manually though, so this shouldn't be an issue anymore. just set ADMIN_USERNAME and ADMIN_PASSWORD env vars and it'll autocreate the user and collection on first run.
https://github.com/ArchiveBox/ArchiveBox/wiki/Configuration#...
You can learn about the origin story / motivation here:
https://github.com/ArchiveBox/ArchiveBox#background--motivat...
https://2020.pycon.co/en/talks/5/ (a conference talk I gave about it)