For this project I wanted a consistent file format for my entire collection.
I have a bunch of stuff I want to save which is behind paywalls/logins/clickthroughs that are tricky for wget to reach. I know I can hand wget a cookies file, but that’s mildly fiddly. I save those pages as Safari webarchive files, and then they can drop in alongside the files I’ve collected programatically. Then I can deal with all my saved pages as a homogeneous set, rather than being split into two formats.
Plus I couldn't find anybody who'd done this, and it was fun :D
This is only for personal stuff where I know I'll be using Safari/macOS for the foreseeable future. I don't envisage using this for anything professional, or a shared archive -- you're right that a less proprietary format would be better in those contexts. I think I'm in a bit of a niche here.
(I'm honestly surprised this is on the front page; I didn't think anybody else would be that interested.)
2/ Proprietary format: it is, but before I started I did some experiments to see what's actually inside. It's a binary plist and I can recover all the underlying HTML/CSS/JS files with Python, so I'm not totally hosed if Safari goes away.
Notes on that here: https://alexwlchan.net/til/2024/whats-inside-safari-webarchi...
'Save the webpage as I see it in my browser' remains a surprisingly annoying and fiddly problem, especially programmatically, so the niche is probably a little roomier than you might initially suspect.
You may be interested in SingleFile[1]
[1] https://github.com/gildas-lormeau/SingleFile
I use it all the time to archive webpages, and I imagine it wouldn't be hard to throw together a script to use FireFox's headless mode in combination with SingleFile to selfhost a clone of the wayback machine.
It's just "Firefox".
https://github.com/webrecorder
It has a standardized format and acts like a recorder for what you see.
Seems JS developer just want their code to the obfuscated and unachievable as possible unless it is via their web server.
These SPA bundles are on the order of megabytes, not kilobytes. You want your users, for their own sake and yours, to be able to cache as much as possible instead of delivering a unique megablob payload for every page they hit.
Is it really? I remember hacking around with with JavaScript's XMLSerializer (I think) like 5 years ago and solved that for ~90% of the websites I tried to archive. It'd save the DOM as-is when executed.
Internet Archive/ArchiveTeam also worked on that particular problem for a very long time, and are mostly successful as far as I can tell.
There's a lot of interest in 'digital preservation' and perhaps one sign of how it's very much early days of the field - it's tricky to 'just save' the results of one of the most basic current computer interactions - looking at a web page.
If you need more than that, then WARC is probably the best. For my measly needs of just preserving exactly what I see, serializing the DOM and saving the result seems to do just fine.
It's not that there aren't workarounds, it's that they are clunky and 'you can't actually save the most common computery entity you deal with' is just a strange state of affairs we've somehow Stockholmed ourselves to.
One category that the archivers do poorly with is news articles where a pop-up renders on page load which then requires client-side JS execution to dismiss the pop-up.
Sometimes it is easily circumvented by manual DOM manipulation, but that's hardly a bulletproof solution. And it feels automateable.
Ultimately, this is the best viewer I've found so far: https://replayweb.page/
It looks like there's a lot of tools for creating them, but not a a lot for viewing.
What they really need is browser support, or at least an extension so a browser can open the files directly.
That's probably the wrong thing. What browsers really need is a thin but standardized API that lets any third-party app that the user has installed on their machine to supply the content for various fetch/reads.
You'd open the WARC in Firefox or Safari or whatever, but Safari et al wouldn't have any special understanding of the format. It would know that your app does WARCs, though, and then knock on the door and say, "Please tell me the content I should be showing here; I'll defer to you for any further "requests" associated with the file/page loaded in this tab—just tell me the content I should use for those, too."
One of the main use cases for an archived web page would be to share archives, and in that case I think you'd want them to be double-clickable with little fuss.
> Although Safari is only maintained by Apple, the Safari webarchive format can be read by non-Apple tools – it’s a binary property list that stores the raw bytes of the original files. I’m comfortable that I’ll be able to open these archives for a while, even if Safari unexpectedly goes away.