Forking Chrome to turn HTML into SVG
fathy.fr
fathy.fr
I'm very surprised to hear this. So printing, either to PDF or to actual printers, may reveal more information about what was drawn to the canvas than normal display, especially if no effort has been made to remove overdrawn paint records. That can have an interesting, if only hypothetical, consequence...
I just hope the code around password entry fields is carefully audited. That's all on the client.
Edit: I found this but I'm not sure it's the one I'm remembering: https://www.techdirt.com/2014/01/28/new-york-times-suffers-r...
Stupid autocorrect
As a simple example imagine that an image is drawn to the canvas and then blacked out. You wouldn't expect that the saved PDF may contain those as separate layers.
Of course this highlights an existing issue with complex formats. You need to be very careful before sharing complex documents.
This is actually a somewhat common method when it comes to a bit of corporate sleuthing.. anytime you see a pretty website with vector-y graphics, maybe engineering-drawing representations.. if the data hasn't been stripped completely or redrawn you can extract information that otherwise people would assume unknowable.
In a recent example... I did this on a startup company's page involving a product where they had a CAD-like side view drawing of one of their products... but the base file (in this case it was an SVG) driving the page actually contained multiple hidden views of the same product and other products and at the 'real' precision of what likely was a DXF export from a CAD program, given to the web team. This allowed a critical dimension of an unannounced product to be precisely determined (to three significant figures) which was a spec that had not been publicly released...
This is actually really interesting. Do you get to do this often?
Also within PDFs (and svgs) you normally clip the area you’re going to draw into to bound it (sort of like overflow:hidden) and anything outside of that doesn’t display, but it’s still there and accessible.
I marvel more at the fact that software is capable of figuring out all the occlusions so you can print the stuff on a plotter. Cad drawings have up to 2M individual vectors in them. Its impressive that it works at all to be honest.
> I'm very surprised to hear this. So printing, either to PDF or to actual printers, may reveal more information about what was drawn to the canvas than normal display, especially if no effort has been made to remove overdrawn paint records. That can have an interesting, if only hypothetical, consequence...
The canvas API is all imperative code, so you might think it’s fairly opaque. That’s what I thought anyway, until recently I hacked on a someone’s generated art demo, mostly to glean insights into the algorithms used. As I looked through the actual canvas-specific code, it struck me that 1) it’s exceedingly statically analyzable and 2) that the imperative APIs could trivially translate to an incremental SVG rendering, because their primitives are nearly identical apart from the imperative/declarative distinction.
Mentioning this mainly because if there’s anything interesting to learn about a particular usage of canvas, it would probably not be a huge investment to learn it. Either by static analysis or by rendering canvas calls incrementally to SVG, anything overdrawn or obscured is sitting right there to inspect without any special browser-internal faculties.
<!DOCTYPE html>
<canvas id="tutorial" width="150" height="150"></canvas>
<script>
var c = document.getElementById("tutorial");
var ctx = c.getContext("2d");
ctx.beginPath();
ctx.moveTo(75, 50);
ctx.lineTo(100, 75);
ctx.lineTo(100, 25);
ctx.fill();
</script>
and used Chrome to pdf print it. Then I opened PDF in a Sumatra PDF Reader, zoomed in and it's obviously not vectorized. window.onbeforeprint = () => {
window.canvas.width |= 0; // clear the canvas
draw();
};
So I guess, for now, this behavior only manifests when you use the beforeprint event.[1] https://yari-demos.prod.mdn.mozit.cloud/en-US/docs/Web/API/C...
To be honest, it works quite well, but there are quite a few bugs in chromium's pdf rendering, especially when it comes to determining the correct page width to apply media queries to, which sometimes affects the accuracy of these SVG's.
I have found printing to PDF from Safari instead of Chrome yields better results when I have gone through the same process. It probably depends on the source material though.
If I remember correctly, the text was split into separate objects by Chrome to reproduce kerning offsets.
I'm genuinely curious if there are any advantages in Chrome->SVG as opposed to Chrome->PDF->SVG.
Are there any graphical effects (e.g. produced by CSS, like blurry text shadows or something) that PDF can't render without falling back to bitmap but SVG can?
Or is there other data that SVG usefully preserves that PDF discards, such as actual source text strings used for text? (As opposed to PDF where getting text out, e.g. when copying to clipboard, usually involves a lot of ugly "reverse engineering".)
My advice to everyone re pdfs is to crack them open by running `mutool clean -d file.pdf` And opening in a text editor. They’re just a tree (well, graph, I guess) of obvious objects.
Ps: mutool convert does a good job of converting from pdf to svg in a fairly faithful way.
PDF seems a bit more bonkers because you render text as strings of glyphs and the conversion back to text is an afterthought. There’s a ToUnicode map that says “glyph 8 in the embedded font is an ‘X’” but that’s there for copy pasting / searching - not for rendering. PDFs are built to render glyphs at positions.
Edit: to go full meta, there are Type3 fonts where each glyph itself is defined as a PDF graphics stream. Which actually leads you in to what’s inside a font. Guess what? lots of them look just like PDFs inside, because the glyphs are defined in postscript. Fonts are PDFs kinda grew up together, and once you start digging into them the similarities are striking.
Epub should be the defacto standard when the included data remains important.
If you go direct to SVG the capture will use the screen css and not be paginated.
[0] https://chromedevtools.github.io/devtools-protocol/tot/Emula...
Basically, convert everything to an archival format, then I'll browse the archive instead of whatever adversarial server side / javascript junk the site is serving.
Either way, you'd probably want your proxy to wait to for any onload Javascript to run before snapshotting the page.
Right now our pipeline looks like this:
1. Code generates SVG. It contains quite a lot of elements and takes something like 250 KB unzipped or 20 KB zipped.
2. SVG is converted to big PNG using resvg. Now this SVG is something like 1.5MB unoptimized. We further use pngquant to shrink it to something like 500KB.
3. We use HTML templates to produce a document with embedded PNG image.
4. Now we use HTML to PDF (could be Chromium, but we use some Java library) to produce PDF document and this is the end result.
Good thing about this pipeline is that SVG and HTML are somewhat easy to understand and modify.
This PDF document obviously contains raster and not vector image, so it does not look good when zoomed and it takes more space that I'd like.
What I want to implement: keep HTML part (because that's kind of report and it should be changed if necessary) but embed image with some kind of vector approach, so PDF would contain vectorized image.
I tried to just embed SVG as it is. Well, it kind of worked. But Chrome printed out enormous pdf. Something like 100 MB I think. My computer almost choked trying to render it. But it was vector, yes.
I've given up on this task as I didn't find any easy approach and generating PostScript for the whole document seems not appropriate. I think I could generate PostScript file for the image, but how do I embed PostScript file into HTML so PDF print would use it as it is?
There’s many libraries out there that can convert vector SVG directly to vector PDF. Sometimes it’s worth just paying the license for, say pdftron. Alternatively if you are generating the would-be SVG content in JS you could directly generate PDF instead of SVG with http://pdfkit.org/
Chrome usually does a good job rendering SVG directly to PDF. It’s overkill of course, but if you’re already used to using it then it’s the path of least resistance. If you’re getting giant PDFs output then I suspect there’s something weird with your input SVG. It might contain some pathological content (huge embedded images? paths with large amounts of almost invisible details? fonts that aren’t embedded sensibly? e.g. vectors of every glyph)
I dont really want to turn entire document into svg. There’s some text, some tables, things that fit perfectly for HTML and not so much for SVG.
Right now I’m thinking about using GNU groff. It would be a very different approach.
The MuPDF ‘mutool’ is pretty handy for dumping internal PDF content to see which parts are large.
I do stuff like this (vector representations of the DOM) for taking screenshots. Why?
1. High resolution screenshots are great when you're sharing from a low resolution device, or when you need to scale them up. I've seen enough crappy screenshots of Twitter in YouTube videos to last me the rest of my life.
2. If your device does sub-pixel anti-aliasing, then your screenshots all have noticeable color fringing around their text. The text rendering is done well before the data hits the buffer that the screenshot is capturing. A fun party trick is to identify someone's OS based purely on a screenshot of some text on a webpage.
3. On Linux (and maybe elsewhere, IDK), color correction (e.g. gamut mapping) is done (in X11) before the pixels get to the buffer that you capture. So with most screenshot tools, you end up capturing a bunch of distorted colors which you then have to map back to sRGB if you want them to look right in color calibrated software.
You can frequently get away with printing a PDF and then rendering that out to a large PNG. In some cases, though, figuring out how to set the page size to match what you seen on the screen can be near-impossible, and more importantly in Firefox there's no way to disable print media CSS when printing a PDF. (You can do this in Chromium.) If you need to edit the image afterwards or want to put it on a website or something, this is far easier to do with the SVG format than with PDF.
I run a little API that converts URLs and HTML into PNG/PDF/SVG.. MP4 too, and the quoted part resonates :)
I recently started delving into the chromium src code in order to try and figure out the reason why max-width media queries don't seem to trigger at the expected viewport/page width, when printing to PDF, but it is quite the rabbit hole.
When I saw html2svg the first thing I wondered was whether it would have the same issue as printing to PDF.
I’m wondering about them as alternatives to frame capture for remote browsers isolation to save bandwidth.[0]
Also related, the Chrome Debugging Protocol exposes a similar bit of info in the LayerTree domain: you can actually get the canvas draw instructions to render a webpage on a canvas.[1]
[1] https://chromedevtools.github.io/devtools-protocol/tot/Layer...
[0] https://chromedevtools.github.io/devtools-protocol/tot/Layer...
For your use-case I'd recommend doing something similar but using the SkPicture structure instead. It would cut the conversion overhead, should support every Skia features, and with some tuning it would allow you to efficiently split bitmaps from vectors (allowing you to send the bitmaps once, and only vector changes after that).
Something I like about using Skia for this use-case is that it allows for zero-latency scrolling.
I believe PDF.js incorporated some form canvas2svg to try and get a SVG backend working which would allow high resolution printing to PDF but not sure where that's at. I believe printing through PDF.js is blurry due to memory constraints since with normal canvas pdf pages just end up as bitmaps sent to the printer.
SVG ends up staying vector through Chromiums print pipeline resulting in much less memory usage while having much higher dpi final output. I would imagine this is due to SVG being turned into Skia drawing commands that end up as PDF that then gets printed through PDFium?
That's curious.
Anyone know why?
Flutter currently has 2 ways to run something on the web: 1. CanvasKit. Primarily, this uses webgl. Though, the app has to download a kind of webGl runtime on the first launch, iirc. If the browser does not support openGl, it will use Skia with a Canvas frontend, leading to blurry and poor performance results 2. webRender. This is flutter's way of trying to make a HTML DOM, but its not that great either. It's inconsistent with the rest of the flutter implementations, and has performance issues because it's not really mature/optimized and has a virtual Dom.
I think an exciting use case would be something like 1. Instead of the blurry image and bad performance of canvas redrawing, it might try to manipulate an SVG in the browser. This is pure speculation tho, correct me if I'm wrong.
How do I best follow your project and any upcoming announcements?
That looks like a Web Archive (WARC) and a PNG screen shot. I think you can make a screenshot with CasperJS. The WARC can be created by wget.
Source?
Wikipedia says: https://en.wikipedia.org/wiki/Archive.today#:~:text=Web%20pa....
So I assumed they don't use it but idk for sure.
for example i made this one[1] with tailwind but i just ended up taking a png screenshot
[1] https://github.com/sentriz/socr/blob/master/.github/socr.png
I'm currently building a web-app for building web pages, and I'd love for the user to be able to view a thumbnail gallery of all the pages they've built. This tool would allow me to build a zooming feature pretty easily.
Outside of those two, I'd imagine the use cases are fairly limited.
Imagine flying through your site in 3D (or even VR) with full control over timing, being able to explode and un-explode your DOM elements as they transition into being - the type of thing that only Apple would do for their WWDC demos with dedicated visualization teams.
The start is to be able to see the rendering engine as a generator for not just raster data over time, but vector data over time. Of course, there's a lot of work to do from there, but this is the core leap.
Here's the thing: this "core leap" you mention isn't new. It's been made long before, on input side: that's what HTML is. All those use cases you mention should be possible, but aren't. Why? Because the for-profit web doesn't want that.
Most websites on the Internet today exist not to be useful, but to use you. For that, it's most important that the website author has maximum control over what the users see. This allows them to effectively place ads, do A/B tests for maximum manipulation (er, "engagement"), force a specific experience on you, expose you to right upsells in the right places. If they could get away with serving you clickable JPEGs, they would absolutely do that. Alas, HTML + JS + CSS is the industry standard, the all things considered cheapest option - so most vendors instead just resort to going out of their way[0] to force their sites to render in specific ways, and supplement it with anti-adblock scripts/nags, randomizing DOM properties, pushing mobile apps, etc.
To be fair, they do have some point. Look at this very thread: currently, the top comments talk about using Chrome -> SVG pipe to access accidentally leaked commercial and government data, such as hidden layers in product CAD drawings, or improperly censored text[1]. Your own example, "export someone's interaction session with a site, pixel-perfect, into DaVinci Resolve or Blender or Unity" is going to be mostly used adversarially (e.g. by competitors). My own immediate application would be removal of ads, which presumably have distinct pattern in such rendering, as they get inserted at a different stage of the render pipeline than the content itself.
This is just the usual battle over control of the UX of a website. The vendor wants to wear me down with their obnoxious, bullshit UX[2], serve me ads, and use DRM to force me to play by their rules. I want my browser to be my user agent. We can't have it both ways[3], and unfortunately, Google is on the side of money.
--
[0] - Well, to be fair, a lot of this is done by default by frameworks, or encoded in webdev "best practices".
[1] - Obligatory reminder: the only fool-proof way of publishing partially censored documents or images is to censor them digitally, then print out, scan back, and distribute the scan. If you don't go through analog, you risk accidentally leaking censored information or relevant metadata.
[2] - Like e.g. every single e-commerce platform. The vendor hopes I'll get tired and make a suboptimal choice. I want to pull vendor's data into a database and run SQL queries on it, so I can make near-optimal purchase decisions in fraction of the time. This "core leap" you mention would be a big win for me, which is why it won't last.
[3] - At this point, accessibility is the only thing that's keeping websites somewhat sane. There's plenty of apologists for all the underhanded and malicious techniques that are core to webdev these days - but they can't usually dismiss the complaint that the website is not usable on a screen reader. For some sites, it would be illegal to do so.
I'm quite familiar with anti-scraping and anti-ad-blocking countermeasures, and the first thing any such tool would block is a non-standard rendering engine like this - so unless the website creator consents, this really doesn't hurt or help consumer-friendliness (which, I agree, is in a sorry state these days) in any meaningful way.
[0] https://frankgroeneveld.nl/2021/08/24/most-underused-browser...
In my case I actually render them on the client using dom-to-image[2] and then store/cache them using R2. This is basically to save on server costs (it's hosted on free CF pages with a worker). A more secure implementation (with a compute budget) might use a headless browser like this server side to render the image.
[1] perfectpollie.au
[2] npmjs.com/dom-to-image
There are countless use cases where a system needs to generate something printable, such as a PDF report, ticket, gift certificate, etc. And generating something that looks great from a system that is good at generating html can be a challenge. Being able to convert a rendered html page to a vector graphic opens up a lot of options.
I already have an AWS Lambda function running Weasy Print for when I need simple PDF generated from a webapp. It'd be great to be able to switch to something like this that does a better job at rendering HTML, and therefore having the option to make more beautiful PDF's.
https://github.com/vercel/satori is another interesting recent project that goes bottoms up to achieve something similar. Def a lot smaller aimed at edge rendering, but at the tradeoff of significantly more random compatibility / rendering issues.
Both approaches have their tradeoffs.
We are improving the DX further so that the feedback loop of what features are supported is nicer, with advice on what to do when some CSS property is unimplemented.
I guess doing too much in a shared context is a security issue.
But I suspect that spelling is on purpose, because it uses the root of the word. If that's the case and there is a story you want to tell, I want to hear it :)