Show HN: CLI tool for saving web pages as a single file
github.com
github.com
How do you guys handle the security aspect of executing stuff like this on your machines?
Skimming the repo it has about a thousand lines of code and a bunch of dependencies with hundreds of sub-dependencies. Do you read all that code and evaluate the reputation of all dependencies?
Do you execute it in a sandboxed environment?
Do you just hope for the best like in the good old times of the C64?
$ git clone https://github.com/Y2Z/monolith.git
$ cd monolith
$ cargo install
These are the ones I used: $ git clone https://github.com/Y2Z/monolith.git
$ cd monolith
$ sudo docker run --rm -w "$(pwd)" -v "$(pwd):$(pwd)" -u "$(id -u):$(id -g)" rust cargo install
That isolated the build process. Similar method to isolate the execution of the built project: $ cd target/release
$ sudo docker run --rm -w "$(pwd)" -v "$(pwd):$(pwd)" -u "$(id -u):$(id -g)" rust ./monolith https://www.grepular.com
Slightly related: I released a project last night for easily running node applications inside containers: https://gitlab.com/mikecardwell/safernode - Without even having node or npm installed on the host system, you can still run commands like "npm install" or "npm start" to run node applications safely isolated inside ephemeral containers.To me, this is kind of like saying you should just run stuff as root, because there might be a privelege escalation vulnerability which lets the code run as root anyway.
Correct me if I'm wrong.
My goal was to make things more secure, not completely secure.
Previously, dodgy libs could read (and add) ssh keys into ~/.ssh/, take over my NPM account by fetching ~/.npmrc, grab a copy of my ~/.bitcoin/wallet.dat, and add a keylogger into my ~/.bashrc
Now, at least it has to break out of docker first.
But I never said it was preferable to run directly on the host. There are other choices.
> My goal was to make things more secure, not completely secure.
There is no such thing as completely secure. The argument against docker is more along the lines of "is it really as secure as people think it is?"
> I've heared this too, but as far as I know it's only because there are potential bugs in the container software that allow the malware to escape.
I'm not sure docker was designed for the purpose of secure isolation, so if it fails to securely isolate, I'm not sure it would count as a bug.
The cgroups interfaces don't offer much security stuff directly -- they're mainly about containing groups of process within certain resource consumption quotas, and afaik, don't really attempt to contemplate secure isolation directly.
LXD approaches this by adding a uid/gid translation layer, so that the uid/gid for anything within an unprivileged container will be offset by a specified value, e.g., calls with user ID 1000 in a container are made to present to the host as user ID 1000000. This comes with its own host of issues which LXD tries to hide.
The short answer is that if security is any type of priority for the system in question and you want to run containerized processes, you should use an OS that implements container security directly in the kernel, like FreeBSD with jails or illumos with Zones, instead of depending on getting exactly the right configuration between all the moving pieces in the Linux container stack.
I could counter your "VM is better than container", with "Separate hardware is better than VM".
I very much like this quote from Alan Kay:
"It doesn't matter what the computer can do, if it can't be learned by billions of people."
There is no good technical reason why modern operating systems can't work out some some scheme for sanboxing arbitrary programs by default. It is obviously necessary. I imagine something like "applications" folder where every subfolder automatically becomes an isolated "container". It would have to be designed with security as the primary concern, though; unlike current container solutions.
Who’s to say that the user’s intent by running the program they just downloaded, isn’t to—say—overwrite a system folder? (Oh, wait, that’s exactly what Homebrew does, with the user’s full intent behind it!)
There are tons of attempts to do what you’re talking about. Canonical’s “snaps” are a good example. As well, every OS sandboxes legacy apps by default (because they’re already virtualizing them, and sandboxing something in a virtualization layer is easy.)
But none of those solutions really work for the “neat FOSS hack script someone wrote” workflow we’re talking about here, where you build programs from source and run them for their intentional side-effects on your system.
You might suggest that there could be a shared sandbox for all the POSIX-like utilities to interoperate in. But what if you’re attempting to use those utilities against your real documents? (For example, a bulk metadata auto-tagging and auto-renaming utility, to get TV episodes from torrents loaded into Plex correctly.) How do you draw the line of what such a program can operate on? AFAICT, you just... can’t. Its whole purpose is to silently automate some task. If it requires constant security prompting, the task isn’t automated.
You copy or move documents inside the specific sandbox.
If you want a pipeline, you establish a chain of inbox/outbox folders.
Obviously, most of this should be done by the OS, not the user.
The workflow:
- You click "download" in your browser.
- When it's done and you click on your download, the OS asks how you want to open it. Instead of "execute" option you get "run in a sandbox" option.
- You type in the name of the sandbox, the app gets copies to /apps/sandboxName or something of that sort.
- The system automatically creates /apps/sandboxName/inbox and /apps/sandboxName/outbox.
- To process a file in some way, you drop it into inbox dir.
For command line, the only change would be switching from "executable pulls arbitrary files" to "I push specific file to the executable".
zip -r squash.zip dir1
becomes | -r squash.zip | zip dir1 |
Start pipeline. Push squash.zip as an argument to zip, get the output. Zip would be the container name.Or, for a simpler, more obvious example: find(1), grep(1), etc. A set of utilities that can all be asked the equivalent of "read literally every file the VFS has access to and tell me whether they match an arbitrary-code-execution predicate." Do you want to literally copy your entire hard disk into the 'inbox' of these utilities, in order to get them to search it for you? (And before you say "well, we can trust the base utilities that ship with the OS to do more than arbitrary third-party utilities"—there's a whole competition of grep(1) replacements, e.g. ag(1), rg(1), etc. Do you want to make it impossible for people to innovate in this space?)
Or how about Nix, or GNU Stow, or, uh, Git? These utilities become useless if they have their own sandbox. Does your git worktree live in Git's sandbox? Vi's sandbox? The inability to make this distinction functional is why mobile OSes only have fullly-integrated IDEs!
Or how about shells themselves! (Or, equivalently, any scripting runtime, e.g. Ruby, Python, etc.) Should people not be allowed to install these from third parties?
Or, the most based example of all: make(1) [and its spiritual descendants], and the GNU autotools built atop it. How does ./configure work if you can't detect true properties of the target system, only of the sandbox you're in?
Well, let's think about the goal here. grep reads files and outputs lines from those files. It needs full read access to everything you want to search. It does not need write access outside of its sandbox. It does not need direct access to network sockets, audio stuff and so on.
Is it unreasonable to create a readonly "view" of the filesystem inside grep's folder? Is it unreasonable to have "files" representing network access, microphone, audio? It will have visual representation in file manager without the need to create custom UI. It could be manipulated by drag-and-drop OR command line. More importantly: it's easy for users to understand. "This app lives in a box. You can put things in that box for the app to use."
>Does your git worktree live in Git's sandbox?
Yes? I mean, I currently have a folder called projects. All my git stuff is in there anyway.
>Vi's sandbox?
If you want multiple sandboxes to be able to operate on a directory, you create "views" for that directory (readonly or read/write) in multiple sandboxes. This shouldn't be some sort of mind-bending idea, considering Unix has symlinks, hardlinks, and mounted filesystems of all sorts.
>Or how about shells themselves! (Or, equivalently, any scripting runtime, e.g. Ruby, Python, etc.) Should people not be allowed to install these from third parties?
There is no reason why a Ruby executable should have unlimited access to the entire file system. Especially if you're only using it for a specific purpose, like serving a website.
What I'm describing here isn't some novel, mind-blowing idea. It's simply dependency injection. With file-based user interface. Every single part of this had been done in various operating systems or programming environments more than once. It's just a matter of combining it all in a sensible way.
You have SElinux for that, if you like bureaucracy and filing triplicate forms to able to run scripts with side effects.
Care to elaborate?
It's okay to do things that don't scale at all, and also okay to make a proof of concept and let someone else worry about scaling it up.
That said, you might want to look at Chromebooks and the Windows and Mac app stores to see what's going on with containerization beyond mobile. (Also, web browsers.)
Isn't that basically just Qubes OS? https://www.qubes-os.org/
Snap [1] goes pretty far in this direction. Apps are isolated against each other, and with AppArmor isolated from the system (at least on Ubuntu, your distro might vary). Android does much of the same.
A big problem is that most software exists to manipulate data on the user's machine, so isolating the software from the User folder is impractical. At the same time this data is usually the most valuable thing about the entire computer. That makes it fundamentally very hard to design a system where you can trust arbitrary apps. Android tried to solve this with a "file open" dialog that's controlled by the OS so that there's an easy way to give apps temporary access to single files, but that leads to weird UX.
See OLPC Bitfrost
I've shunned that assuming it'd be a big slow down, but I do keep meaning to at least try it, uh, after I knock it.
(No, containers like anything else aren't and haven't been completely secure all of the time since and forever, but it'd take a more sophisticated - and certainly deliberately malicious - tool to do any damage to your system, or to files you didn't explicitly allow it access to.)
Meanwhile, my desktop remains clean and ready to play media in native environment with good hardware support.
I used to think it'd be slow, until I tried it. My computer is 8+ years old, and it works fine. I mostly do text work.
But, at least the way I was doing it, it's not adding any security as discussed here, since you're doing everything in the VM so anything in the VM has access to everything just as if everything on the host anyway.
What they don't have access to is my actual desktop, nor yesterday's snapshot of the VM desktop.
It's obviously not as fast as running GNU on bare metal, but it's fast enough for text work.
I guess this is where we set the bar in 2019.
I'm not GP commenter, but certainly what I do for a living is browse the the internet and edit text files.
For me, the VM was my entire machine. (It wasn't meant to improve security, it was purely because I wanted Linux but couldn't have it on the host.)
I wanted GNU tools and environment, and I also wanted to not have to install Linux on the Apple hardware, because I don't have the patience for that.
I personally prefer to hope for the worst. This way when nothing happens I feel extra lucky, and if bad things do happen, I feel proud of being ready for it.
I usually do a quick check to see if there's any red flag. Like URLs or Base64 blobs. I also try to stay away from programs written in languages whose environment I don't know so I can check if any dependency stands out wrt what the program claims to do.
> Do you execute it in a sandboxed environment?
Either I run it on an untrusted server, or on my laptop as an unpriviliged user (with no access to X/wayland).
Why of course. I do this for every piece of software on my computer, from the device drivers to the OS, I review every patch to firefox & chrome as well. /s
Running someone else's software inherently means extending them trust. This objection is especially confusing on a piece of software where you can actually inspect all the source if you like (unlike e.g. device drivers and OS code unless you run Gnu/Linux + all Free drivers, as few but RMS do).
But on that last bit I disagree. Many, many systems run nothing but stock open source kernel drivers under Linux. I daresay home systems with closed drivers are more of an exception. All those VMs in the "cloud".
https://www.archive.ece.cmu.edu/~ganger/712.fall02/papers/p7...
You want to run the software properly sandboxed, since linux doesn't have really engineered OS-level solutions (à la Solaris) that means vm e.g. running in qemu, or biting the bullet and switching to qubes.
It's an objective one.
> There has been a massive amount of work put into securing containers for a variety of use cases using both layers available and by adding to the kernel.
That doesn't change the fact that security was never the primary goal for containers, so secure containers were and are a bunch of tricks, kludges and prayers being built up in the hope that eventually all the holes in the model will be patched.
The "Making containers safer"[lwn][hn] talk was literally two days ago. Note how it says safer, not safe.
[lwn] https://lwn.net/SubscriberLink/796700/9bc9daa32a8fe499/
I suppose when security enhancements are made to any other system to make them safer (i.e. everything in the realm of security), you apply the same logic? Subjective.
Do they ever offer to run your container in the same VM as those of other customers?
They never do this. For secure isolation, they only trust VM isolation. It seems unlikely that this will change.
> What would even be the point of user namespacing, network namespaces, filesystem namespaces, etc... if not security?
They're for installation/configuration/administration. They allow you to run multiple applications on one Linux VM, and to configure them independently, almost as if you were running multiple VMs (with the advantage of lower overheads - only one instance of the kernel).
Kubernetes puts this to good use, letting you treat application deployments as commodities across your cluster.
Containers do not offer secure isolation. They are by nature much leakier than the isolation VMs can offer. The Docker folks still treat isolation-failures as bugs, of course. (Well, ignoring things like the way 'uptime' gives the uptime of the underlying machine, and not of your container.)
I don't disagree they are leakier abstractions but they can still satisfy a wide variety of workload security needs.
Docker itself could be called a huge kluge, at least compared to Solaris 'zones' and FreeBSD 'jails'.
They're similar to containers, but are supported directly by the kernel, whereas Docker has to pull together different kernel features to create its abstraction. [0]
> What's to stop some state agency inserting its code into the core? No way to review everything
1. This isn't a point about containers, it's a point about Free and Open Source software in general. Do you avoid all Open Source software when security matters? 2. I'm pretty sure the Linux kernel folks review everything, and I imagine the Docker folks do too 3. You're implicitly assuming that closed-source software is safe from government pressure. It is not.
[0] https://blog.jessfraz.com/post/containers-zones-jails-vms/
We don't. Companies that produce proprietary code are not immune from attacks on their repository, and are more vulnerable to, say, bribery. They're also more vulnerable to attacks on their distributed binaries - users do not have the option to compile from source, so you compromise every user this way.
Proprietary software is also far more likely to embed 'telemetry' spying, or to use sloppy security practices and rely on security-by-obscurity. Authors of Free and Open Source software know that they (generally at least [0]) cannot get away with this kind of thing.
It simply isn't true that proprietary software is more trustworthy than FOSS. If anything, the opposite appears to be true.
AUR, Arch User Repository.
> thus you are expected to vet the packages yourself
Obviously, as a long time Arch user I didn't know this.
It is the Arch User Repository.
> DISCLAIMER: AUR packages are user produced content. Any use of the provided files is at your own risk.[0]
What about javascript execution ? If you replay your capture, you have no idea of what you will see on general Web2 website.
The only way I know to capture a web page properly is to "execute" it on a browser.
Gildas, the guy behind SingleFile (https://github.com/gildas-lormeau/SingleFile) is well aware of that and his approach realy works everytime.
Try on a Facebook post, a Tweet, ... It just works.
Tbh, often those are superfluous, or egregious examples of bad web dev, so it seems a reasonable solution for most cases.
SingleFile is a different approach, but it's a lot more involved/less convenient than a cli, and loading in something like WebDriver on the cli for this would be overkill, unless you're doing very serious archival work.
When you create a modern webapp, a lot of data are retrieved from servers as Json and formated in the browser in Javascript. Even sometimes Css is generated on browser-side. Even more, on webapp where user login is taken into account, the display is modified accordingly.
That's the web of 2019. The approach consisting of geting remote files and launching them in a browser is really naive.
Speaking of SingleFile, it as a cli version and can handle full web 2.0 webapp without any problem. And of course, the Web 1.0 webapps work as well.
In terms of React at least, fetch requests are not a part of the framework in any way and any present would typically be done in custom code in lifecycle methods. Even Redux, is—by default—client-side only. Stores are in-memory, actions populating them would make fetch requests with React/Redux-independent logic.
Other JS frameworks are, typically, the same. And all of that is just considering dynamic XHR. Loading scripts is much less typical, and never required. The most common application of this I've seen is the GA snippet, which mainly does it to ensure the load is async without relying on developer implementation: it's 100% unnecessary to do it this way.
So yes, unless you're distributing a tracking snippet that you expect non-devs to be blindly pasting into their wordpress panels and still have it work efficiently, generally speaking use of this method is never necessary, and commonly a red flag for poor architecture.
Web 2.0 refers to the use of ajax. This refers to the early 2000s sajax, jquery..
If you want to separate angular, react, vue maybe it's web 3.0.. but wasn't web 3.0 referred to as mobile?
Ajax will* still work fine with an internet connection as long as those ajax endpoints don't require cookies and don't linkrot.
* not 100% sure how the tool handles relative URLs embedded in source : if it's not clever enough though, this is very fixable via PR (as in its not an architectural limitation)
The main problem with too many websites is they've become too much about technology and have left visitors, as well as the spirit and intent of the internet behind.
I agree with you. BadSite: "we want you to experience this.."
GoodSite: "we want you to learn this..."
But for run-of-the-mill websites? Such tools are not only overkill, they break the spirit of the internet.
If it does capture the tracking shite, would disconnecting from the internet be a good enough blocker?
* mthml encodes the page as a multipart MIME message (using multipart/related), essentially an email (you're usually able to open them by replacing the .mth by .eml)
* WARC is its own thing with its own spec
* WAFF is a zipfile, not sure about the specifics
* webarchive is a binary plist, not sure about the specifics either
Your tool generates straight HTML which any browser should be able to open. It probably has more limitations, but it doesn't require dedicated client / viewer support.
Maybe once you've got all the fetching and extracting and linking nailed down it would be a nice extension to add "output filters", but that seems more like a secondary long-term goal, especially as those archive formats are usually semi-proprietary and get dropped as fast as they get created (WARC might be the most long-lived as it descends from the Internet Archive's ARC, is an ISO standard and is recognised as a proper archival format by various national libraries).
E.g.
test.maff
`-- 1566561512/
|-- index.rdf
|-- index.html
`-- index_files/
`-- ???
When I was messing around with archiving things locally I settled on WAFF, because it's pretty much trivial to create and to use. Even if your browser does not support it, you just need to unpack it to a tempdir and open the index file.Tbh. I had arrived at the conclusion that Mime would be great, but it never struck me that someone had already made a "standard" of mime and HTML.
[1] https://github.com/gildas-lormeau/SingleFile
Edit: This question just came to mind. If MHTML saves images using base64, and base64 dataurl images have a limit size, how would you save extremely large photos? Take for example the cover image of this article https://story.californiasunday.com/gone-paradise-fire. When I saved the page in MHTML format, the re-rendered image showed up quite blurry compared to the original. Was the size limit the cause?
[2] https://stackoverflow.com/questions/12637395/what-is-the-siz...
Many years ago used a program that could convert them but the name escapes me. A brief search shows a few results that appear to do a similar conversion, potentially they may be of use.
> This question just came to mind. If MHTML saves images using base64, and base64 dataurl images have a limit size, how would you save extremely large photos?
From what I've read there's no official limit for base64 encodings though IE/Edge limit them to 4GB according to caniuse.com. Haven't personally encountered an image or GIF that was too large not to be saved in an MHT (I have over 10k MHTML files saved, some image and GIF heavy ones up to 200MB each).
Also a sibling comment corrected me about the data URI use, MHTML uses a separate scheme but nevertheless still uses base64 for encoding.
For that example article you linked it seems likely to be the way the program you're using is handling the Javascript-loaded images. I saved it in Vivaldi (Blink-based engine) and the main image displayed at full res when opened locally while the other images didn't, while when saved with a pre-Quantum Firefox using the UnMHT addon it saved all the images at their fully loaded resolutions. Some MHTML saving implementations clearly have advantages over others it would seem.
There no needs to insert MHTML page into IFrame!
If you need insert MHTML content into IFrame, just convert it to HTML+JS firstly.
Does it? IIRC it stores assets as MIME attachments, hence the "M": the result is not HTML (which this would I assume be), it's a multipart MIME message whose root is an HTML document.
edit: in fact when downloading mht files osx / safari misrecognises them as exported emails and appends the "eml" extension.
[1] https://techdows.com/2019/06/google-removes-save-page-as-mht...
- how do you handle images?
- does it handle embedded videos?
- does it handle JS? to what extent?
- does it handle lazily loaded assets (i.e. images that load only when you scroll down, or JS that loads 3 seconds later after the page is loaded)
In general, how does this work? The current readme doesn't do a decent job explaining what the tool exactly is. For all I can tell, it probably just takes a screenshot of the page, encodes as base64 into the html and shows it.
This tool was able to capture three.js applications and other interactive sites quite well.
SPA's often (but not always) do this. Content is loaded in via React components and such...
The desktop version saves an HTML file, stylesheet and images/fonts locally, and it only contains the HTML of the snippet with the CSS rules that apply to the DOM subtree of the element you select.
I'm still working out bugs but it would be great if people try it out and let me know how it goes.
https://stackoverflow.com/questions/10266334/add-on-to-copy-...
I'm hoping my tool will be better so it's good enough people would be willing to pay for it, but we'll just have to see.
I'm glad there's more people taking a look at the use case, and I'd be interested to see a list of similar solutions.
If you combine this with Chrome's headless mode, you can prerender many pages that use JavaScript to perform the initial render, and then once you're done send it to one of these tools that inlines all the resources as data URLs.
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome ./site/index.md.html --headless --dump-dom --virtual-time-budget=400
The result is that you get pages that load very fast and are a single HTML file with all resources embedded. Allowing the page to prerender before inlining will also allow you to more easily strip all the JavaScript in many cases for pages that aren't highly interactive once rendered.Also, side-question about Rust: how do I get rid of absolute file paths in the executable to avoid information leakage? I feel like I partially figured this out at some point, but I forget.
And about Rust: I think you're way ahead of me here as well, this is my first Rust program. If you're talking about it embedding some debug info into the binary which may include things like /home/dataflow then perhaps there's a compiler option for cargo or a way to strip the binary after it's compiled. ¯\_(ツ)_/¯ Sorry, that's the best I can tell at the moment.
Hope this answers your question (it gets converted to a data URI, and there's apparently no de-duplication).
Need to find all articles relating to 'widget'?
$ ls -l ~/PDFArchive/ | grep -i widget
This has proven so valuable, time and again .. there is a great joy in not having to maintain bookmarks, and in being able to copy the whole directory content to other machines for processing/reference .. and then there's the whole pdf->text situation, which has its thorns truly (some website content is buried in masses of ad-noise), but also has huge advantage - there's a lot of data to be mined from 50,000 PDF files ..Therefore, I'd quite like to know, what does monolith have to offer over this method? I can imagine that its useful to have all the scripting content packaged up and bundled into a single .html file - but does it still work/run? (This can be either a pro or a con in my opinion..)
- Allow for recording source and author information (PDF ... doesn't always provide this).
- Allows for full-text search.
- Allows for editing out annoyances.
I'll frequently go from HTML to some simplified representation (e.g., Markdown), and then re-generate formats that are useful elsewhere: HTML, PDF, ePub, etc.
Dumping from HTML to Markdown frequently makes cruft-removal far simpler, and the principle content of most pages is text. In rare instances, images are useful, and even more rarely, any multimedia content (video, audio, programmatic content).
What's depressing is the number of sites which screw with even basic HTML. E.g., the NY Times rarely use HTML tables for tabular representation, and instead use a homebrew combination of custom markup, CSS, and JS to much the same effect. Pretty, in situ, but brittle and transports exceedingly poorly.
That's just one of many such cases.
JS can be removed from the final document using the -j flag. HTML Files can also be grepped for content, unlike PDFs.
Imagine a paper in the form of a single HTML file, which has (a subset of) the data included, the graphs zoomable, the colors chanegable (to whatever vision problems you have) - maybe even the algorithm to play around with!
Jupyter Notebooks already go in that way. only without the single-file, open in browser aspect, I think.
The firefox extension seems to do that :
[1] https://www.gnu.org/licenses/license-list.html#PublicDomain
One thing I would propose to add - either a flag, or by default - have it parse the path to the page and create the file with the name - that way you can just "monolith {url}" and not have to worry about it.
I am also curious as to how it handles advertisements and google tracking and such; some way to strip out just those scripts (and elements) could be handy.
How does this work on neverending webpages/forever scroll? How will it behave if you need to authenticate before browsing the page?
It seems to work for basic pages quite well, I think that lazy load will work for most pages as long as the JavaScript is embedded (no -j flag provided) and the Internet connection is on. It saves what's there when the page is loaded, the rest is a gamble since every website implements infinite scroll differently.
Authentication is another tricky part -- it's different for every browser. I will try to convert it into a web extension of sorts, so that pages could be saved directly from the browser while the user is authenticated.
Whenever I want to download a video, using YouTube-dl, from a site that requires authentication, I first login using my browser and then exports the cookies using an extension.
I am also interested in cloning/forking sites for modification purposes, I will feedback you on the results four my consulting gigs
It will evolve into a reliable tool in a couple weeks and it should eventually work for embedding everything, including things like web fonts and @url()'s within CSS. If anything doesn't work, please open an issue, I have plenty of time to work on it.
If the outcomes are realistic, take a massive list of sites, make a snapshot of each page, replace the POST login URLs with the phishers, deploy these individual HTML files, and spread the links through email.
I wonder how does this project handle forms.
upd: Done, now forms get their action="/submit" converted into action="https://website.com/submit" when the page is saved.
I suspect for saving videos, a good approach would be some sort of proxy + headless browser combination, where the proxy is responsible for saving a copy of all data the browser requests for.
Thoughts?
I use youtube-dl for youtube and other popular web services myself. Embedding a video source as a data URL could in theory work, but it'd be quite a long base64 line. Also, editing .html files with tens or hundreds of megabytes of base64 in them would perhaps be less than convenient.
HTMLD, WARC, MTHML, MAFF and webarchive are all "container" formats which bundle assets next to the HTML using various methods (resp. bundle, custom, multipart MIME, zip and binary plist).
https://webrecorder.io/ solves that problem by recording all interactions and then replaying them as needed.
> Webrecorder takes a new approach to web archiving by “recording” network traffic and processes within the browser while the user interacts with a web page. Unlike conventional crawl-based web archiving methods, this allows even intricate websites, such as those with embedded media, complex Javascript, user-specific content and interactions, and other dynamic elements, to be captured and faithfully restaged.
It seems to me that having one file that any browser can easily open (and not require Internet connection to view) is a big advantage over having a directory with assets alongside the .html file. It may be one of those things that make things easier yet nobody really complains about how things are usually done when the page gets saved. I hope more browsers add support for saving pages as MHTML in the nearest future so that we wouldn't need tools like this one.
Bookmarks, or downloads, externally.
I'm at 252 posts, presently, just checked. That seems to be a complete log.
After how many entries before the HN software started tripping on your favorites?
On question: How does it handle those cookie pop-ups, gdpr-warnings etc?
That's an interesting question. I think it depends on how the given modal is implemented, but closing them should technically work (unless the page is saved with JavaScript code removed [-j flag]). Those notifications can easily be removed from the saved file using any text editor, should be pretty easy if you know how to edit HTML code. I don't think removing it would violate anything since "this website" will no longer really be a website but rather a local document at that point.