BreezyPDF Lite: HTML to PDF generation as a Service
github.com
github.com
I really don't like that this uses something non-standard for header/footer and stuff in the margins since that's already covered by standardized CSS @page stuff; https://www.w3.org/TR/css-gcpm-3/. I'm using that with Weasyprint for automatic header/footers with auto page-numbering, including setting strings from the html to use for heading names, document title, author, etc... example CSS: `h1 { string-set: h1-title content(), h2-title ""; }` + `@page { @bottom-left {content: string(h1-title) string(h2-title);}`.
Weasyprint seems closest to princexml and is free.
https://docs.breezypdf.com/metadata#use-css-for-page-size
Headers/and footers are just HTML strings and can be super rich with images etc and customized with CSS. Page numbering is free as well in headers/footers.
Of course you could just use properly positioned <header> and <footer> tags and do whatever you need to with JS for page numbering.
Tasked with solving the same problems again, I'd probably look at headless Chromium or Firefox. Their JS PDF renderers are fast and if you're doing enough to keep an instance loaded all the time there's no start up time.
Into a pdf like this: https://www.dropbox.com/s/v4j4n1cvtm032w9/breezy-pdf-dashboa...
Headless browser rendering is fine if all you need is a two-page invoice PDF, but it falls down when you need control of anything other than basic stuff like the font size.
This is definitely not the case anymore, and BreezyPDFLite supports most of the features you mention, while supporting the same HTML/CSS/JS you might be displaying to end users across evergreen browsers.
h2 {flow: static(header);}
@page :right {
margin-bottom: 1.4cm;
@top-left {content: flow(header);}
@bottom-right {content: counter(page);}
}
Edit: Also, tables of contents, with page numbers and links in the PDF: ul.toc a::after {content: leader('.') target-counter(attr(href), page);}Building TOC is just as easy as building the HTML and linking to the ID's appropriately.
Page numbers are supported in header/footer templates, or via manual computation when you render the HTML or with JS.
I'm currently involved in an effort to do the reverse and there isn't a day that I don't curse the PDF specification and the various implementations. And with the 'data:' source for graphical element and MathJax there isn't much reason for for instance scientific papers to be published as pdfs to begin with.
https://github.com/thomaspark/pubcss
PubCSS has the right idea.
We use EvoPDF for this purpose, which also uses a website based webbrowser under the covers. Unfortunately it is quite slow, especially when Javascript is required for the report. It also handles tables badly across multiple pages, and full page backgrounds are also cumbersome.
That it is a supposedly open format is a joke, there is so much old cruft in there that you could implement if you had access to Adobe's source code but unfortunately nobody but Adobe does.
PostScript (which PDF is a restricted version of in many ways) is clean, but PDF is not.
`#{chrome_alias} --headless --disable-gpu --print-to-pdf="#{pdf_path}" "#{html_url}"`
This uses Google Chrome headless for the actual PDF generation.
We now switch to a Microsoft Word rendering backend where we process Word files with template strings in them and then run Word in headless mode to save files as PDFs. While the HTML-to-PDF approach works, most of our users work with Word all day so we are solving the wrong problem.
Clever, but that does sound a bit messy.