ArchiveBox/ArchiveBox: open-source self-hosted web archiving
github.com
github.com
What I do right now "to collect, save, and view sites you want to preserve offline" is by use of a Firefox plugin called WebScrapBook. Click-click-done, and I have a local searchable (!) copy of a webpage exactly as it looked in the browser. With styles and all, in one file. WebScrapBook is pretty highly configurable.
In the future I would like to have a solution that doesn't require some Firefox plugin.
The recommended install includes a search engine that works well, aside from a few false positives due to being fuzzy search. I don't have much in it yet so I can't say what the performance is like once you reach e.g. thousands of pages, but I imagine it would still perform well except maybe for mass operations like rebuilding the entire index.
Looks like ArchiveBox has more export options? EDIT: looks like ArchiveBox is focused on continuous change tracking rather than than just snapshots like Wallabag.
Key benefit for me is having actual local files. The resulting PDFs are searchable on their own, so I can sync those back to my Mac for reference (and Spotlight indexing). But the HTML snapshots are also pretty decent.
One thing I’ll be looking into is automatic tagging (since it’s a Django app there are plenty of likely ways to inject that info).
Their roadmap is also very interesting: "v2.0 Federated or distributed archiving + paid hosted service offering"
https://github.com/ArchiveBox/ArchiveBox/wiki/Roadmap#v20-fe...
One neat tiny implementation detail of ArchiveBox that I just highlighted on HN today is our use of asymptotic progress bars when we don't know how long archiving a page is going to take: https://news.ycombinator.com/item?id=27860022
How can I make archivebox ignore such errors and continue with the rest of the websites?
Command used:
archivebox add < exported_bookmarks.html --depth=1Also note you got the arguments backwards, make sure to put the file redirect at the end, after --depth=1.
It's also possible that was the last URL in your list and it stopped at the end. If you want to re-try archiving you should run this instead:
archivebox updateI know other people would still have that tab open from 3 weeks ago, but I just don't work like that.
I'm not going to complain about wayback/archive.org at all, but the nature of the beast is that there's certain requests they have to obey - and with my own offline, non-exposed equivalent, I don't (well, I do, but I simply don't receive them)
I create browser bookmarks regularly but I wouldn't be bothered to SSH into my server to also tell it to grab a copy of the URL. Automating this with a browser plugin would be cool.