It's real next gen thinking on this topic.
As for the featured tool wayback... If HN readers can't figure out what it does after reading docs, its likely the thinking behind it is equally unclear.
It's real next gen thinking on this topic.
As for the featured tool wayback... If HN readers can't figure out what it does after reading docs, its likely the thinking behind it is equally unclear.
Keeping it simple, I download pages in Markdown adding some metadata (some tags). When I want images or more I use singlefile extension. Add Recoll to the mix and that's all I need.
https://github.com/deathau/markdownload
https://github.com/gildas-lormeau/SingleFile
https://www.lesbonscomptes.com/recoll/pages/index-recoll.htm...
I loved this project when it was called 22120 and was here https://github.com/c9fe/22120
It was AGPL3 and it was great! http://web.archive.org/web/20210123003545/https://github.com... The author links to an interview with an open source focus publication about it!
I should not have allowed the credit of the previous clearly explained open source project to carry forward after renaming, because it is now a different project.
I am sorry. I apologize. I will not recommend this software again, and I will try not to make this mistake in the future.
Thank you for helping me see this error.
There's a long-existing web archiving ecosystem with established formats for recording and publishing archives.
The repo linked could probably do with some clarity itself in how it does or doesn't fit in with standards.
But it's not open source and pretty limited regarding the use case.
The thing is, archiving is a multi-faceted and hard problem (i.e. video content, live streams, interactive sites), and will remain so. A complicated task leads to complicated tools.