Wkhtmltopdf: Command line tools to render HTML into PDF
wkhtmltopdf.org
wkhtmltopdf.org
In the meantime Chrome added headless support and with command line or puppeteer you can do the same thing with a much newer and more secure engine.
However if you need full CSS-Paged support then PrinceXML is probably your only option and its not cheap: https://www.princexml.com
There are reliability problems with wkhtmltopdf, but it's possible to make it reasonably reliable with enough error detection and retrying.
So, it would become fairly inaccurate using wkhtmltopdf.
It would look like either the entire item had no issues, or it had sustained damage everywhere.
Basically prince adds a watermark on the first page - so you do all your revisions and then only pay for the final copy to remove the watermark (however as any editor would know, we always have many final, final final, final v3 versions ;) so paying for prince once-off at 500$ is probably the best way for anyone who needs professional pdfing.
I wrote a latex / markdown / html editor with a simple mechanism (using the browsers default print settings), I think it’s not too bad for those looking to pdf from html:
Issue with mine is you would have to do all style sheets in line. I would recommend serious users to either pay for Prince if they want header/footer/page number control or just use print style sheets and code the html / css themselves and use headless chrome or chrome itself to print. Many people want JavaScript in their pages, I can’t recall if headless chrome can correctly account for that. File > print is the best imo :)
Yep, this is the way to go now. Very simple API with puppeteer and everything looks the same as when viewing the page in a browser (save any @media print css).
I already tried to fill in a bug: https://bugs.chromium.org/p/chromium/issues/detail?id=953313... but the spec is not very clear nor requires the most compact arrangement possible. It would probably be a worthy addition to the spec even just for the environmental impact (limiting number of pages when printing).
I like the functionality from Chrome to use it headless, but my main problem with it is that text is not selectable in the rendered PDF. It seems to be an screencapture as image or something. Anyone has this working? Selectable text is really needed for generating PDF invoices.
Here is reproducible demo that is selectable and clickable (links):
- https://majkinetor.github.io/mm-docs-template/docs.pdf
This is done using puppetear.
The same can be achieved by Save as PDF in chrome of single page containing entire site:
- https://majkinetor.github.io/mm-docs-template/print_page/
This is all part of my mm-docs framework:
I just tried printing the Hacker News front page to PDF in Edge which is the same engine as Chrome and the text was selectable.
Source: working on automatic PDF reporting in my current company.
But in the end, I switched[2] to "chrome --headless --print-to-pdf-no-header", since it reproduces browser behavior pretty much by definition and, while it's a colossal dependency, it's also trivial for non-technical users to install.
[0] https://github.com/mikepqr/resume.md [1] https://weasyprint.org/ [2] https://github.com/mikepqr/resume.md/commit/206a6cbc85fd0456...
I ended up swapping in Google's Puppeteer library to render my PDFs, and despite needing a bit of plumbing to get my Scala build script talking to the node.js runtime (and I needed to do e.g. page numbers and table-of-contents extraction myself using Apache PDFBox) in the end it worked much better. Things looked the same in puppeteer PDF as they did in the browser, which is something I could never quite achieve with wkhtmltopdf
Many docker containers exist for making Latex (the whole texlive[1] distro) into a service.
[1]: https://www.tug.org/texlive/doc/texlive-en/texlive-en.pdf
Although HTML/CSS/JS aren't fully viable for print just yet, they're not missing much.
I think HTML/CSS is viable for print (JS has no business in print as print is not interactive), as it is used a lot. The point I'm trying to make is that Latex is simply better equipped due to it being specifically designed for page layouts.
How long do we have these H2P tools? And how far have they come in terms of supporting page layouts?
And the same for bare pictures since the rendering engine has to have drawn all of that to display it?
Is it something to do with pdf's being Adobe's?
(I still see the utility of a command line program but just wondering since they have some issues around font-sizing, fonts loaded through the web, etc)
Considering there are competing pdf readers to Adobe, this makes sense
Browser pdf styling is quite good actually, I just hate the fact it prints date / web titles by default (only user can unselect) and doesn’t work that well with paged media
I was thinking of building my own project in an adjacent area, so I did a lot of research. PrinceXML is my favourite, although it doesn’t work well as an integration to an app since then you have to share profits with them
In the end I ditched my project (lol)
It is but I remember going through so many hoops to do it, when a "browser.export_pdf()" would have been so handy.
Does yours live in some repository online?
https://github.com/puppeteer/puppeteer/blob/v5.3.1/docs/api....
It is a CLI tool written in python and quite usable in my experience.
[0]: https://itextpdf.com/en/products/itext-7/convert-html-css-to... [1]: https://github.com/flyingsaucerproject/flyingsaucer
It makes it super easy for front end devs to layout and generate PDFs using tools they already know.
Can't beat the price either, especially compared to coldfusion or some of the .net libraries out there.
the top half of this line on page 1, the bottom half on page 2.
I didn't find a way to CSS out of that. To be fair, my customer never allocated budget to find a solution so that is probably not very important to them in our scenario.
Not fixed yet.
I have been experimenting with Mozillas JavaScript library Readability (Powers Firefox's Reader feature) to convert wikipedia articles to cleaned up text. It works quite nice for everything that works with Reader in Firefox or Safari.
$ google-chrome --headless --disable-gpu --print-to-pdf-no-header http://www.example.com/ %.pdf: %.*
pygmentize -f html -O full $< | wkhtmltopdf -T 0 -B 0 --page-width 210mm --page-height 5000mm --no-background - - > .tmp.pdf
pdfcrop .tmp.pdf $@
rm -f .tmp.pdfThe fact that they are achieving it via poly fill sounds like a genius approach!
[1]: https://wiki.contextgarden.net/Installation
[2]: https://dl.contextgarden.net/myway/tas/xhtml.pdf
[3]: https://pandoc.org/
Can this handle svg or JavaScript libraries like Charts.js? I’ll check the details but if you know off the top of your head :)
Thanks again for the link
At some point we'll likely move to chrome/puppeteer but in most of our tests the templates required modifications to mimic the existing PDFs so it's something we've been putting off.
But there are ways to do this same thing with Puppeteer which uses a modern version of headless Chrome. It’s pretty straight forward and there are easy to chase down documents on how to do it.
Does that pose a security concern to you? It might. Are you accepting any data from the outside? Are you requesting anything from the outside world? Does your rendering browser have a place to write any data?
I've been using Apache FOP with some success.
This is more-or-less what we do in our application when we want to present a PDF for review or perform drawing on top of PDFs in contexts where we cannot rely on PDF libraries for whatever reason.