URL to PDF Microservice
github.com
github.com
Just try ?url=file:///etc/passwd on the demo instance.
That seems to be a quite common issue with services like this built on generic libraries.
It might be better to blacklist file:// rather than trying to have a comprehensive whitelist.
window.location.href = 'file:///' will return console error: "Not allowed to load local resource"
{"status":400,"statusText":"Bad Request","errors":[{"field":["url"],"location":"query","messages":["\"url\" must be a valid uri with a scheme matching the http|https pattern"],"types":["string.uriCustomScheme"]}]}[1]: https://url-to-pdf-api.herokuapp.com/api/render?url=https://...
Sometimes you can get rid of the sticky headers with &emulateScreenMedia=false parameter if the page has well implemented @media print rules in CSS. We decided to use page.emulateMedia('screen') with Puppeteer to make PDFs look more like the actual web page by default.
Pages which use lazy loading for images may look incorrect when rendered. &scrollPage=true parameter may help with this. It scrolls the page to the bottom before rendering the PDF.
Using these options make the PDF better: https://url-to-pdf-api.herokuapp.com/api/render?url=https://...
At least with PhantomJS I felt like my system would begin to lockup if there were too many instances rendering at the same time (and it didn't appear to be an issue of too little memory).
Nonetheless, this looks promising.
It is scalable. It scales linearly (and for practical purposes indefinitely) with the amount of money you spend on AWS Lambda. It might not have a nice constant factor, but it is scalable.
No, some things can't be, like badly architected monoliths, or databases. Especially with databases it's not a given, which is why for quite some time in the last years things like the "MongoDB is web scale" blew up and people started mindlessly asking "is it scalable" (which as you figured out, means very little for a lot of systems). I'm also pretty sure that your example scales worse than linearly, since you have to introduce multiple levels of management at some point.
"scalable" = "can it be scaled", which is not a given for every system
Also, for the record, this comment I'm making right now is just pedantic.
Round-robin those folks.
I wonder if this part of chrome could be easily extracted as a C++ library.
The PDF-ization isn't the part that's hard to extract (there are libraries to create PDFs from scratch already available, and they're small/intelligible). Rather, it's the rendering of a webpage for display that's the hard part, and what most of the code in any web browser is concerned with. Whether that display is a monitor or a PDF doesn't change much.
There's also room for improvement in how efficiently a single server instance can render PDFs. The API doesn't yet support resource pooling, this would make reusing the same Chrome process (with e.g. 4 tabs) possible. The implementation requires careful consideration since in that model it's possible to accidentally share content from previous requests to the new requesters.
Is that the limiting factor? How many would you optimally do in parallel if RAM wasn't an issue?
Link to size bug: https://github.com/GoogleChrome/puppeteer/issues/666
Headless Chrome is quite new so it still has some bugs, but I have a hunch that it will in the end have most reliable and expected render results.
I've written somewhat extensively about the deployment approaches here: https://hackernoon.com/more-than-you-want-to-know-about-head...
https://service.prerender.cloud/screenshot/https://google.co...
Thanks for packaging this up!
One of the small things I've recently let myself be bothered by is how divergent HTML/browser snapshots from web.archive.org, archive.is, and Google cache can be, even for relatively simple pages. I've already given up on trying to make (not even sure if it's a good idea) HTML look nice as PDFs.
Edit: this seems to be a take on saving the HTML plus assets - https://github.com/pirate/bookmark-archiver.
Edit: ahh someone else mentioned it too.
One of the biggest values I think this yet-another-PDF-service has is that if you open the print preview on a desktop Chrome, it should be really close to what the API renders. Should make debugging a bit easier.
The main use case is to render content generated by yourself, e.g. receipts and invoices, but I don't see a reason why it couldn't be used for rendering news or blog articles.
[1] https://chrome.google.com/webstore/detail/singlefile/mpiodij...
Not sure of more straight forward hosting options
Show HN: Kozmos – A Personal Library | https://news.ycombinator.com/item?id=14980075 (Aug 2017)
specifically: https://addons.mozilla.org/en-US/firefox/addon/scrapbook-x/ and https://chrome.google.com/webstore/detail/worldbrain-the-res...
https://github.com/pirate/bookmark-archiver python script
https://chrome.google.com/webstore/detail/singlefile/mpiodij... save as a single html file
> uses "data URI" scheme to embed image and frame contents into the page : the resulting format is not MHT/MHTML
PDF generation is 'fun', and I've tried several of the different options out there for HTML to PDF generation, but settled on wkhtmltopdf.
E.g. on a multi-page invoice, show a sub-total row at the bottom of each page. Does anyone know how to create this kind of function?
The browser will display a different CSS to the printer, so to speak.
I can't use an externally hosted service like this because some of my URLs are non-public. So when the user requests a PDF, I render the HTML to a temp file on the server, invoke chrome via command line, and serve up the converted PDF.
Right now, I'm using wkhtmltopdf which supports custom headers and footers and I'm mulling over how to do the same with this solution.
Stuff like receipts, etc to be converted to PDF will contain customers' information - not to be put on a site with no info on how it's used.