Show HN: Pdf-Bot, an API/CLI for Generating PDFs Using Headless Chrome
github.com
github.com
https://github.com/GoogleChrome/puppeteer/blob/master/exampl...
Really easy to work around, I used it to build a simple CLI to generate device screenshots of a webpage by modifying the user-agent and resolution to match each device.
I've had a great experience working with Headless Chrome to convert webpages to PDF for my side project EmailThis (https://www.emailthis.me).
It uses Puppeteer by Chrome DevTools team - https://github.com/GoogleChrome/puppeteer.
If you try saving a any discussion website (HN, Stackoverflow, Reddit), you will get the PDF.
Please file a bug report if you think this is an important feature.
Yes, CMYK support is very tempting feature for the whole print industry.
However, I think this is necessary if you want to fit it into a microservice with a REST interface. For REST, I think the usual expectation is that a) the request returns quickly and b) you can submit any number of requests in parallel. Given that loading a page into headless chrome, rendering it and generating a pdf is both resource intensive and time consuming, I guess you need some way to decouple that process from the interface.
https://developer.mozilla.org/en-US/docs/Web/CSS/page-break-...
One of the best ways I've used to generate PDFs is by using a DOCX as a template, and replacing certain placeholders within the document (a DOCX is a ZIP containing a few XML files). It's great if you work in a corporate environment, as it's easy for non-technical staff to make it look exactly how they want and it's easy to update (just replace a file and check it works). You can use headless LibreOffice to convert DOCX to PDF.
Depending on your use case, I feel like you guys might be interested in http://weasyprint.org/. It is an open source HTML to PDF converter written in Python. It passes the Acid2 test and implements CSS Paged Media.
I recently discovered https://github.com/ArthurHub/HTML-Renderer formerly know as https://htmlrenderer.codeplex.com/
PDF generation [...] 100% managed (C#), High performance HTML Rendering library
I am currently using athenapdf[1] but I will have a play with pdf-bot.
I had a quick look at `pdf-bot`, and though we both rely on the same underlying technology (we are only just moving to headless Chromium; we were on Electron before), I believe we have slightly different ambitions with our respective project. But, I may be biased.
For example, `pdf-bot` seems to be tied exclusively to a specific converter, and storage backend. With `athenapdf` however, we are moving more, and more towards building a toolkit or rather, framework for other people to construct their own conversion processes (or even microservice)[0].
Consequently, we are working towards general abstractions like fetching, converting, and uploading, that can have different implementations (e.g. wkhtmltopdf, LibreOffice, Weasyprint, etc).
With our microservice assembly as well, we are focused heavily on ensuring we have:
1. Instrumentation, and metrics (which `pdf-bot` appears to currently lack)
2. Support for different retry mechanisms (e.g. retry using the same converter or retry using a different converter)
3. Support for multiple input MIME types
4. Synchronous API calls (`pdf-bot` appears to be mostly asynchronous, with batch processing, and callbacks)
5. Ease of installation (e.g. Docker), and configuration
We also have a CLI assembly[1] that can support custom JavaScript plugins[2] (e.g. Markdown -> PDF, Readability, etc). So you don't need to run a service or make API calls for conversions.
[0] https://github.com/arachnys/athenapdf/tree/v3/pkg
[1] https://github.com/arachnys/athenapdf/blob/v3/cmd/cli/main.g...
[2] https://github.com/arachnys/athenapdf/tree/v3/pkg/runner/plu...
My only small problem with it was the somewhat complex setup for using athenapdf-service with a new project (especially since I use docker-machine) but I have now mostly automated the whole thing.
Just out of interest - do you consider asynchronous an advantage (being a Node developer I generally love async very much)? Not that it matters to me - my needs are trivial for the service to handle.
Edit: actually I can see how it async would make my life much more complicated for my simple use case - I would have to write something to track requests and responses rather than just looping through a bunch of URL's that need converting.
We actually went with Docker for the set up because it simplified dependency management tremendously, and it allowed us to deploy on platforms like Kubernetes, Swarm, and ECS. As a plus, it gave us some confidence that if it works for us, it should work for others (obviously, we have come across cases where Docker behaves differently across platforms).
I consider asynchronous processing (in this context) as advantageous in some cases. Indeed, when we were refactoring `athenapdf`, we considered introducing a message queue for workers to pull work from, and to put back when the work completes. The problem with this however, is that we can't as easily scale horizontally (i.e. introduce node replicas behind a load balancer), as if we tried to get / update a job, we may not get the same node we originally got. I mean, the solution can be as easy as introducing a centralised message queue of sorts (or even a sticky session), but that complicates the set up process, so we decided against it.
Taken together, for our specific use cases, we believe it is a lot simpler to consume a synchronous API. No webhooks / callbacks. No polling. No concerns over acknowledgement. If a HTTP call fails, we will know about it immediately. If a complex retry mechanism is needed, we think this should be accomplished in the client application.
In the long term, I believe we should have a toolkit that can easily be plugged into a wider orchestration engine like Conductor (https://netflix.github.io/conductor/). That way, anyone can develop their own conversion process pipeline with ease.
I was wondering, does it support custom page headers (or generally running elements)
Currently, the only print ready HTML to PDF processor that I know is Prince [1] and to a lesser extent Firefox.
This other project[2] was recently on HN but you can't find it easily on github search.
[1] https://github.com/search?q=topic%3Aheadless-chromium&type=R...
Mozilla's documentation (https://developer.mozilla.org/en-US/Firefox/Headless_mode) is still incomplete.
I don't know if there's an easier way or service these days
These days I use Calibre instead because the overall experience is better, but for single articles it's a bit overkill.
[0] https://github.com/michaelrsweet/htmldoc [1] http://user.it.uu.se/~jan/html2ps.html
A better comparison would be against the likes of wkhtmltopdf[0], which uses webkit, or the pdf generation features of phantomjs.
1,047 Open 975 Closed
Yep, that's about how I remember it. It was such a pain to build on Windows (especially to get a single static binary) that people contributing fixes would often attain hero status by attaching a random binary to an issue. Specifically, GIF support was broken on the official Windows build for 4+ years:
https://web.archive.org/web/20140917181225/http://code.googl...
Page.printToPDF: https://chromedevtools.github.io/devtools-protocol/tot/Page/...
There are many "free as in beer" (closed-source), freemium, and/or free trial options offered as a carrot leading to a commercial product. Most have a watermark and/or page count limitations.
http://selectpdf.com/community-edition/ (5 pages max)
(With no disrespect, of course, to the authors of these libraries - they just didn't work well for us.)
ebook-convert file.html file.pdf --pdf-add-toc