SingleFileZ, a web extension for saving pages as HTML/ZIP hybrid files
github.com
github.com
It's a strange format, but I think it falls in the "good enough" category. Still I never see it in the wild
WARC is also standarized by ISO and has a nice spec that's pretty easy to understand[2]
@gildas I saw there was a comparison table[3] but it seems to be missing WARC. Could you shed some light on why?
- [0] https://en.wikipedia.org/wiki/Web_ARChive
- [1] http://digitalia.sbn.it/article/view/1473
- [2] https://iipc.github.io/warc-specifications/specifications/wa...
- [3] https://github.com/gildas-lormeau/SingleFile#file-format-com...
With that said, I do understand your motivation.
- [0] https://www.archiveteam.org/index.php?title=The_WARC_Ecosyst...
So you've definitely seen it in the wild, perhaps without knowing about it.. about every email is one.
Assuming I download webpages from via ssl/TLS, would there be a way to also save their criptographic signature so that the resulting file, along with the website certificate, could be verifiable, possibly in court?
I've seen a number of situation where malicious public clerk do not update ab official public website with information about upcoming events, and then updating it once it's too late.
I'm wondering whether I could use Https features to bring such actors to court.
I saved this page with both webscrapbook and singlefilez. Both archives looked the same. Webscrapbook's was 22.4kb while singlefilez was 65.9kb. I unzipped singlefilez and rezipped it with higher compression and got it to 20kb but it wouldn't open in the browser.
While the size doesn't really matter, what I don't like is that singlefilez renamed the images to sequentially numbered files (1.gif, 2.giff, etc.) and css to stylesheet_0.css while webscrapbook kept the original names of the files. I would much rather it kept the original file names.
The additional 40KB corresponds to the part which self-extracts the zip file (in order to view the page without installing any extension). Note that the original URLs can be found in the comments of each entry in the zip file.
It can take full HTML files and internalize the CSS + HTML and compiles them into a .zip file.
The major difference is that its an Electron app and not a web extension though we're about 80% of the way done porting all of it to a web extension.
In retrospect, I would have done this as an EPUB.
Our users have asked for other features like taking multiple pages and building them as one 'book' and this feature is supported in EPUB.
Also, it would mean Polar would support EPUB natively anyway which is another big feature we need as we only support PDFs right now.
I lot of people here mention MHTML and WebArchive formats.
I think my main criticism of these is that EPUB has more universal 'reader' support and EPUB 3.0 is basically just HTML in an enclosure anyway.
What would you do if your HTML extraction script had a bug in the pre-compiled form? I guess you're just stuck?
I guess that's not the end of the world.
The EPUB form in a future version of Polar would at least require an EPUB reader which makes it a bit heavier.
In addition, the backup is not as usual the naive backup of files but a real backup of the page as interpreted by the browser.
For me it is the best component for saving web pages from far away.
The SingleFile saved content is saved in Zip format in the result Html page.
The final SingleFileZ html page has Html header, Zip file in the body and a bunch of Javascript to decompress the content when the browser opens it.
It's magic !
https://addons.mozilla.org/en-US/firefox/addon/scrapbookq/
Mostly because I used to use the older “Scrapbook” add-on (before it stopped working in Firefox 60) and I still have a number of pages saved in that format – ScrapbookQ is, with some effort, compatible with those saved pages.
As in take a website I am developing and have my server serve SingleFileZ instead of what I would usually.
On paper, you could even store an entire website in a SingleFileZ file. It just need to be implemented...
FYI, when I saved github readme page animated gif location was blank.
Thanks.
I use it regularly, and it creates a single pure HTML file with inline medias. Easy to read anywhere.
So appart from the compression gain, why the need for SingleFileZ ?