SingleFile: Save a complete web page into a single HTML file
github.com
github.com
I mean, if I want to save pages over the next 11 months, should I install SinglePage or SinglePage-lite?
SingleFile was ultra-valuable for this.
If anyone has a similar use-case, I wrote some pretty rough (and slow) code to post-process SingleFile's output to remove any HTML that wasn't contributing to the presentational render by launching puppeteer and comparing pixels. It's available here: https://github.com/mieko/trailcap
I like these because it makes it easier for me to make manual edits when necessary and it's a better solution for long term archiving (IMO). But I would love to add your project to my workflow.
Now, having said that, the text in SingleFile-Lite's "Notable features of SingleFile Lite" sound like a list of issues :-P. It looks like these are issues with Chrome, but do you know if/how these "improvements" will affect Firefox?
- ScreenToGif to record video sequences and produce the final GIF: https://www.screentogif.com/
- Macro Recorder to record and replay user navigation: https://www.macrorecorder.com/
- Blender to edit the video, add text comments, and make the intro: https://www.blender.org/
Taking a snapshot of my user's screen and then display it to them later (maybe in an iFrame)?
For example, try saving my home page: https://andrewrondeau.herokuapp.com/
The img tags are converted correctly, but there's still <figure class=image><a href="https://andrewrondeau.herokuapp.com/... in the single HTML file.
Even better, just search for ="http
Edit: I think I now understand the issue. I confirm SingleFile saves only the current page and not linked images, for example.
I'm going to hijack your post for a question! I love the way you can use the editor and select "format for better readability," then save just the stripped down version of the page. I use this to send it to my e-ink device.
The question I have is whether it's possible to toggle the default save to use the formatted version automatically? I dug into the options and didn't turn anything up!
- Annotation editor > default mode > edit the page
- Annotation editor > annotate the page before saving
Really!? You don't think it's trying to be condescending to try to find criticism on someone's praise, haha? Heh, aanyway :) You feel product is bad!? So weird!! I guess you find there what you bring to it. Wonder what you're protecting there, if you share more of your thinking, we can do know you more. Even so, I think we can just celebrate gildas' achievement! :)
It's too complicated but at least it works!
Well that goes to show its longevity I guess.
[1]: https://opensource.apple.com/source/CF/CF-550/CFBinaryPList....
I have Waterfox-Classic and unMHT (fished out of the Classic Addons Archive, just remember to turn off Waterfox's multiprocess feature) since I occasionally need to archive web pages - and more importantly, reopen them later.
mhtml is just MIME, literally every discrete URL as a MIME part with its origin in a Content-Location header, all wrapped in a multipart container. I don't understand why it's not a default format.
Quantum was the the project name to re-engineer Firefox internals, with lots of design changes, not just plugins. XUL/XPCOM APIs were dropped, as an occasional programmer I understand why, "Quantum broke my plugins" is a reasonable first approximation for most users.
Lemme see if I can pull up the command I use to mirror doc sites.
wget \
--recursive \
--level=5 \
--convert-links \
--page-requisites \
--wait=1 \
--random-wait \
--timestamping \
--no-parent \
$1I thought MHTML was NOT standardized which is why it wasn’t across all browsers yet. From what I remember, every company was doing their own implementation of it. Maybe it’s gotten more standardized the last few years though.
The alternative format (used by the Internet Archive and Wayback Machine) is WARC. It's also a single file, but it's preserving the HTTP headers as well; so its applications is specifically for archival purposes. [1] The "wget" tool which is co-maintained by the Web Archive people also has support for it via CLI flags.
Though when it comes to mobile browser support I'd recommend to use MHTML, because webkit and chromium both have support for it upstream.
Though, on MacOS, WebKit tries to migrate most APIs to the Core Foundation Framework, which makes it kind of impossible to implement as a non-Apple-employee because it's basically a dump-it-and-never-care Open Source approach. [3]
Don't know about chromium (my knowledge is ~2012ish about their architecture, and pre-Blink).
[1] https://github.com/WebKit/WebKit/tree/main/Source/WebKit/Net...
[2] https://github.com/WebKit/WebKit/tree/main/Source/WebKit/Net...
I mean, in principal curl would run on the other platforms, too...but as far as I can tell there's an initiative to move as much as possible to the CF framework (strings, memory allocation, https and tls, sockets etc) and away from the cross-platform implementations.
I would also download entire websites using a software which name I forgot, to read them offline. Back when websites held in a single floppy disk.
Good times!
And GitHub supports embedded videos in README.md files, videos are generally smaller than GIF files and their disabled autoplay is a feature = you save your data until you press play.
Using data so wastefully like this always reeks of privilege to me - especially on something like GitHub. Wikipedia, for instance, never allows things like this.
Because it's a relatively new feature, and probably, a lot of devs don't know about it (I didn't).
I did this [animated gif] once actually, before the feature was introduced, and I definitely hated it, but I had no choice.
Thanks for bringing this to the general attention, though :)
If the demo sequence is <5 seconds, I have never found myself becoming impatient. Gif is perfect for very brief demos. Anything longer than that and I'd like to have some idea where I am at in the video stream (and other controls as indicated)
Yes, the gif bothered me too :D
True since May 2021 so I think a lot of people are still finding this out...
In my experience GIF is still the most set-it-and-forget-it way to know a video will play, to get cross-platform support out of mp4 you may have to provide two different codecs. Anyway, not disagreeing with you and most gifs could drop 90% of their size with better choice of resolution and framerate. This readme is particularly egregious doing a screen capture with scrolling.
As for saving bandwidth until you want to play, I haven't tried this yet but it seems adequately clever to wrap a loading=lazy gif inside a details/summary tag: https://css-tricks.com/pause-gif-details-summary/
From the wiki page for Quick Sync [2]:
> Intel Quick Sync Video is Intel's brand for its dedicated video encoding and decoding hardware core. Quick Sync was introduced with the Sandy Bridge CPU microarchitecture on 9 January 2011 and has been found on the die of Intel CPUs ever since.
I can't confirm but I'd guess your performance issues lie elsewhere than in the h264 decoding specifically.
[1] - https://ark.intel.com/content/www/us/en/ark/products/85212/i...
Thanks for pointing that out. I've looked at this table before and payed attention to HEVC, not AVC, so I believe that's where my mistake came from.
[1] https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video#Hardwar...
Accelerated video decode is often disabled by default on Linux versions of browsers and can be quite dependent on versions of drivers/mesa/X-vs-Wayland/etc.
H.264, even on the high profile is not CPU intensive on a 2014 machine. Unless you are watching 1080P with 5-10Mbps, which is not the norm for internet video.
Video codecs are not my area of expertise. Which codecs are these and what tool(s) would you typically use to ensure you provide them?
[0] https://gist.github.com/ingramchen/e2af352bf8b40bb88890fba4f...
Any documentation on this? Because I have tried to embed video in issues and PRs before, and did not manage. I'm hoping such documentation will explain how this extends to issues and PRs.
I think I basically get the idea, what kind of database are you using? Recoll sounds like a good idea, but I'm also thinking about how I might also make this public-ish.
(i.e. I teach in college and would love to have a centralized way to store and search all my assigned readings, which are most often webpages)
Each html page is processed by (1) getting url, title, time saved (this is under-rated as approximate time of saving is useful if you want to rediscover) and then (2) taking a screenshot and finally (3) extracting text with readability.js and hopefully doing some keyword analysis.
Right now it is stored in a local SQLite Database, although the article content is stored in text files. For search, I can use ripgrep to look through the associated text files.
The eventual goal is to create a flask app which will allow for interactive management of the bookmarks (tagging, searching). I've already got static generation of bookmarks.
Here's a screenshot: https://imgur.com/5YP4sP5
Then add a search engine over that, for 'what was that article about long term effects of DDT on ecosystems I was reading a long time ago?' queries.
And you get a memex - a way to outsource part of the brain to a computer :).
BTW, it's not a landing page. This is the README on GitHub.
While writing this comment I found that it lived on as a (now "legacy") new extension named ScrapBook X [1], and then yet another one named WebScrapBook [2], which seems to still be alive!
[0]: http://www.xuldev.org/scrapbook/
[1]: https://github.com/danny0838/firefox-scrapbook
[2]: https://addons.mozilla.org/en-US/firefox/addon/webscrapbook/
> Benefits of the Manifest V3
> - None
wget --mirror \ --convert-links \ --html-extension \ --wait=2 \ -o log \ https://example.com
So one empty space followed by indented text (2 or more spaces)That is not only suboptimal, it is stressing on the server. At least you added a --wait=2, but on any large site/hoster/CDN, this might still get your IP banned or throttled. And on e.g. the English wikipedia this will then take 149 days. Which means that by the time you hit the last page, the first ones (and their links) are out of date.
I'm building a bookmark app, and I plan to use this to save bookmarks!
I'm a simple man, nothing too fancy. Here's a crude demo in progress - https://zewallet.netlify.app/ Follow progress here - https://twitter.com/recursiveSwings/status/14917723874649088...
Would love to have ANY tips or feedback!
I'm definitely in the market for a bookmark service that archives my bookmarks, Diigo stopped working a year or two ago, and Pinboard can't stay up
Add SingleFile? This extension will have permission to:
- Access your data for all websites
- Input data to the clipboard
- Extend developer tools to access your data in open tabs
- Download files and read and modify the browser's download history
- Access browser tabs
... all of which I don't mind, as long as the extension can't exfiltrate any of the data it can access (send it to a third party). i.e. - no network connections from the extension
- no modifying web pages or executing any code in the context of a web page
Does anyone know whether these sorts of extensions can exfiltrate data? -- this is a concern if the project author's credentials are stolen by a threat actor, which has happened before.I'd like to use SingleFile and have no reason at all to distrust it, but I'd like to understand the security impact of installing lots of web extensions. How do people handle security risks like that? Do you run a separate vanilla browser with no extensions for sensitive tasks?
I wish this behavior was more well known and encouraged by Google.
I don’t do this myself, I try to research any extension I add and don’t do automatic upgrades. I use as little extensions as possible.
That being said, it's a bit weird that this kind of tool is even necessary at all. I would have expected native saving to include CSS background graphics as well, but apparently they don't for some reason, so I think this is pretty useful. Until now, I have also used pandoc (--standalone) to merge all resources into a single HTML file which worked great.
https://wicg.github.io/webpackage/draft-yasskin-wpack-bundle...
Here's the bug: According the HTML spec, elements like <h2> and <div> cannot be inside <a> tags. But using js you _can_ push <div>s instead of <a>s. (It happens from document.insert-type functions, frameworks like Angular/React allow this)
Look at nasa.gov, there's html:
<a href="/press-release/nasa-invites-media-to-next-spacex-commercial-crew-space-station-launch-0" date="Wed Mar 02 2022 10:35:00 GMT-0800 (Pacific Standard Time)" id="ember196" class="card ubernode cards--card cards--2row cards--2col nodeid-477815 ember-view"><div class="bg-card-canvas" style="background-image: url(/sites/default/files/styles/2x2_cardfeed/public/thumbnails/image/51846702013_a0cc55100a_k.jpeg);">
<!----> <h2 class="headline"> ...
</h2>
</div>
</a>
After running this through SingleFile you can visually see the changes, but the html changes are: <a href="/press-release/nasa-invites-media-to-next-spacex-commercial-crew-space-station-launch-0" date="Wed Mar 02 2022 10:35:00 GMT-0800 (Pacific Standard Time)" id="ember196" class="card ubernode cards--card cards--2row cards--2col nodeid-477815 ember-view"></a>
<div class="bg-card-canvas" style="background-image: url(/sites/default/files/styles/2x2_cardfeed/public/thumbnails/image/51846702013_a0cc55100a_k.jpeg);">
<h2 class="headline"> ...</h2>
The way that sites like Wayback Machine handle this is by using the web-replay library Wombat https://github.com/webrecorder/wombat that also uses JS to insert those elements.But what the hell! I was working on a similar html-downloading/reproducing tool and this bug really bothers me. I'd either like the HTML reading standard to be updated to accept <div> inside of <a>, or also make that impossible to do via JS.
Interesting. What is it about those pages that makes saving them raise security issues?
(Also what’s up Andrew! YC S09 represent :wave_emoji:)
(I used both and ended up favoring monolith, but can’t remember why. I think they’re pretty comparable/am grateful for both of them)
Sure, it uses libraries to do the heavy lifting, but these are all popular, well-tested libraries with well-scoped feature sets (html5ever for parsing HTML, url for parsing urls, etc).
If you're looking for a tool like this but think you might need to tweak it, you should give monolith a try.
I followed the other advice on this thread. In the options:
- Annotation editor > default mode > format the page
- Annotation editor > annotate the page before saving
It automatically format the page into reader mode then I can click "Save the page" icon to save it. But sometimes I want to download the page as is. Like this thread for example. "Restore all removed elements" button doesn't seem to work to revert the changes.
For now I just set default as mode as normal and enable "annotate the page before saving", and then click "Format the page for better readability" when needed.
I think I'm not the only one who wants an alternative to pocket. A bookmark manager that can archive the links to prevent linkrot.
https://github.com/sergiotapia/ekeko
This is awesome! I would love to integrate this somehow into my project to "singlefile" bookmarks as people make them.
@gildas do you have any recommendation on how to approach this with your extension? Could I run a headless chrome and trigger this extension?
[1] https://github.com/gildas-lormeau/SingleFile/tree/master/cli
[2] https://github.com/gildas-lormeau/SingleFile/blob/master/cli...
Webpage saving technology does not seem to have kept pace with the evolution of the web.
Images loaded by CSS aren't saved at all. JavaScript on the page will often hijack a saved page and not let it display at all.
One option that works fairly well and does not require installing a browser extension is to save the page as a PDF.
I wish browser developers would put more effort in this area.
p.s. the demo is nice
- Read and change all your data on all websites
- Modify data you copy and paste
- Manage your downloads
Is there a way to use a version that requires less of these permissions? e.g. it seems we can address the first permission by only activating it on click, but I'm not sure if that addresses the other ones.
There are many similar tools there, from archiving to rendering.
New browsers don't seem to do this, the create a separate folder for the assets, which is super annoying.
Maybe I'll make a CLI implementation (sorta like wget but with this tacked on...)
For some reason, I went in expecting to see a JS-enabled multi-page web site into a SPA in a single HTML file, but I didn't expect to see images get embedded.
Perhaps offer a recursive traversal option too, but don't try that on Wikipedia :)
> approx
The only thing I miss I wish it was easier to script.
ArchiveBox does indeed look fantastic. Their homepage alone is beautiful.
I bookmarked both ArchiveBox and now also SingleFile, but WebScrapBook gets the job done (in almost all cases).
It's amazing how much CSS is useless in a page. It's especially annoying for SingleFile if it contains images... That's why SingleFile removes (almost) all unused rules, selectors and CSS properties by calculating the CSS cascade.
[1] https://chrome.google.com/webstore/detail/print-edit-we/olnb... [2] https://chrome.google.com/webstore/detail/save-page-we/dhhpe...
[0] https://github.com/gildas-lormeau/SingleFile/blob/15801c8ef4...