Wkhtmltopdf Considered Harmful
blog.rebased.pl
blog.rebased.pl
I built a basic CLI wrapper to make it easy to drop it in in place of wkhtmltopdf: https://www.npmjs.com/package/puppeteer-cli
I've used wkhtmltopdf heavily for years and it is a disaster today. Font rendering is inconsistent across systems, its CSS/HTML support is frozen in time many years ago.
The author's suggested alternative doesn't use a real browser rendering engine, it uses its own CSS/HTML parsing and rendering implementations (!!!). I don't trust it to keep up with standards and don't want content authors to have to deal with yet another dialect of HTML/CSS with its own pile of quirks. We already know Chrome's capabilities and quirks and I'm quite happy with Chrome's print menu and the output of my backend PDF generator being the same thing.
> The CSS layout engine is written in Python, designed for pagination, and meant to be easy to hack on.
Raise your hand if you want to hack on a CSS layout engine in Python while generating your PDFs. I'll wait
After discovering the issues surrounding font embedding and lack of interoperability with windows/adobe acrobat (when redacting) I was tearing my head out trying to figure out how to get it to work. In the end, I chalked it up to being a limitation of how the program handled fonts, and looked for an alternative.
After discovering puppeteer and dropping it in (replacing wkhtmltopdf), basically all the problems I was having with fonts disappearing were solved. Did not look back.
I've been through so many combinations of fop/flyingsaucer/wkhtmltopdf/phantom and pdftk/mcpdf/qpdf over the last decade.... but it seems we're finally about to solve this problem...
The thing makes ImageMagick look well architected and streamlined.
Wkhtmltopdf is far and away the best solution (at least in Ruby land) even if it is often painful to work with.
> TL;DR: replace it with weasyprint.
I, like many people in this thread, have used wkhtmltopdf for years and it works pretty well.
The developers are hard-working and helpful (@ashkulz even commented on a commit on my personal fork!), I just think the job is too big for the resources they have. Too many edge cases.
Also I didn't realize that DocRaptor has an integration with Heroku. I should really look into that.
These are the os dependencies I ended up needing: libxrender1, libxext6, and libfontconfig.
So yeah - the library is fine, but installation ux/docs could use work.
No mention in this of headless chrome - setup as a lambda service is also pretty nice
And their grievance is about abilities to generate PDFs from a template with a predefined presets. Clearly, other libraries will suit that task better.
https://developers.google.com/web/updates/2017/04/headless-c...
Recently, I've been investigating writing new projects in react-pdf
Now we have Chrome headless etc as well.
Just because something is old doesn't mean it's bad. "Old" can also mean "well-understood" and "widely implemented," both of which are true of PDF. I can make a PDF and be reasonably confident that it can be read on any platform under the sun, and that someone needing to transform it will have plenty of tools available to do so. Neither of these things are true of XPS, or really any other competing format save HTML, which isn't really comparable as it aims to solve different problems than PDF does.
There are definitely things to not like about PDF, but its age isn't one of them.
It's not a dig against the format.
Considered harmful is kind of a tired cliche.
It ended up being "8 reasons why I dislike Wkhtmltopdf"