QuestPDF: Modern .NET library for PDF document generation
github.com
github.com
its basically a wrapper for wkhtmltopdf but I develop an app that has probably generated a million +/- invoices/statements over the past 5 years with it, and its been rock solid for me. Was a bit of a bear to get it working the first time (not a ton of documentation that I could find at the time), but once working, was easy to add/change new documents/layouts.
As it uses wkhtmltopdf under to covers, it is a HTML->PDF tool, but I prefer that, at least for my use case.
Not sure there is a dotnet-core version, so that might be a problem for some.
Looking at QuestPDF's API docs, it doesn't look like they support URL / HTML to PDF generation. I think this would be a great addition especially given the age and issues with Rotativa and TuesPechkin on their public repos.
[1] https://github.com/webgio/Rotativa
[2] https://github.com/webgio/Rotativa.AspNetCore
[4] https://github.com/projectkudu/kudu/wiki/Azure-Web-App-sandb...
I don't know if it's due to "partnerships" but I never could understand why Microsoft didn't do better at supporting .NET Word & PDF tooling since .NET Core came out. The older versions I know at least had support for Word docs. Creating documents is a huge foundation of their company.
Some features, like eSignature, there's more co-opetition and partnering.
It makes me nervous having Chrome running on the server, even inside a container without root. Doubly so if the user is able to control any portion of the page being run by Chrome.
Chrome is pretty horrific in terms of memory usage. How are you handling the startup time + memory usage? (with their associated costs)
How the .NET Foundation kerfuffle became a brouhaha
2021-10-08 https://news.ycombinator.com/item?id=28794352#28795511
> the project had now been silently moved to GitHub Enterprise (likely in the short window @dnfadmin had owner access). The author states that projects in GitHub Enterprise can be entirely controlled by the owner of the account (the .NET Foundation). This transfer happened silently.
I had used https://gotenberg.dev/ on AWS in the past. Many of the options available at the time weren't usable in Azure outside of a VM due to needing to make use of GDI interfaces that were disabled for security reasons. Interested to see how it compares to that and the other options being floated at the time like Puppeteer*
I used this as a trick for making Crystal Reports work.
If they understand the same thing with tagged PDF as what is being discussed in this thread, that page says that "Add new APIs to add attributes to document structure node when creating a tagged PDF.", which could be a milestone as old as of 2020 [0]
[0] https://skia.org/docs/user/release/release_notes/#milestone-...
https://github.com/edragoev1/pdfjet
so do the Java, Swift and Go versions.
While the docs are somewhat sparse, Example_45 shows how to create PDF/UA compliant PDF.
And just from looking at the examples you can see how quickly it becomes messy. It looks like something that would work fine for very simple document structures but would get messy really fast once you add any level of complexity. Imagine having to write code for a document with multiple levels of headings, tables and images with captions, parts of text that need emphasis, etc.
From my experience the best solution for generating PDFs in .NET are, as mentioned here before, wkhtmltopdf wrappers. Take a Razor view, add some document properties (page size, margins, headers, footers, etc) in code, and output to a PDF.
I would even argue that on desktop using Office APIs to populate word documents and output them to PDF is more efficient than this.
It all depends on how you structure your code. Of course, you can create huge chunks of HTML+CSS that cover your entire document. Or, to improve maintainability, you can split that HTML into multiple components and compose them together.
Very similar rules can be applied to this library, where you split your code into smaller parts by using properly named methods. That gives you better understandding on the structure and ways to traverse the implementation.
Futhermore, using a programming language gives you many benefits. You can rely on all language features like loops, conditions, methods, formatting, recursion, etc. You have access to IntelliSense, static analysis and refectoring tools.
In terms of changing the layout content - it would be equally difficult when using any markup language. After all, you don't want to parse markup file every time you generate PDF - for performance reasons.
HTML may be a good choice for relatively simple documents, where you don't care that much about splitting content between pages, etc. However, you still will fight with performance problems related to running entire web browser.
Dink2PDF was crashing monthly in this scenario due to internal unmanaged memory problem so we had to replace it. Not to mention, HTML to PDF libraries are insecure, and dink is no exception - you can execute arbitrary code on the server. Not to mention that you need to have full browser engine in your app...
BTW, we used dink for half decade in public web apps with millions of users. Dink was used to create PDF reports and we never had a crash in that concurent scenario. However, when we started doing typical non-web multithreading this started to happen.
[1] https://docs.telerik.com/devtools/document-processing/librar...
1: https://www.w3.org/Style/CSS/Tracker/issues/334
2: https://github.com/w3c/csswg-drafts/issues/4760
3: https://github.com/Kozea/WeasyPrint/issues/93#issuecomment-4...
<?php
$pagenum = 1;
$pagenum++;
print("Page: $pagenum");
?>
These are all things that good page layout software can deal with easily.
HTML to pdf is also pretty slow and unreliable when you want a table of contents or an index with page numbers. On a fast machine, a 200 page PDF can take several minutes to generate. PrinceXML is the only software I’ve tried that does a good job of it. For very simple documents (no CSS, limited Unicode) HTMLDoc is pretty good and very fast.
Content matters for row height, as does fonts, styling, etc. Footers, too. The number of footnotes (thus the height of the footer) can depend on the page content, too.
Maybe there's an image in there, or some sort of specially styled content that makes the line height larger than normal.
You have to be able to render the row to know the final dimensions, to be able to make a call whether it should go on that page or the next. If you don't, then you end up with one page actually rendering over into a second page.
I've heard good things about https://github.com/Kozea/WeasyPrint in Python.
The CSS3 Paged Media spec was born broken on some fundamental things like counter resets, then effectively abandoned in 2013, so some complex print-specific requirements like fully customizable page numbering just don't happen without additional tooling. Accessible tagged PDFs are still a struggle, and I think only Weasyprint readily supports them among free or open-source options (and only since around September).
<?php
$pagenum = 1;
$pagenum++;
print("Page: $pagenum");
?>
As far as pagination is concerned I wrap each page in this tag
<page size="A4"> </page>
with this css:-page[size = 'A4'] { height: 29.7cm; width: 21cm; page-break-before: always; }
page[size = 'A4'][layout = 'portrait'] {
height: 21cm;
width: 29.7cm;
}
So that's the incredibly complex and challenging problem of pagination solved.And fortunately html does layout too.
if (rowcount > 100) newpage();
But lets use an entirely new, esoteric PDF specific layout language instead and we wont call that the "hard part".
Actually to be fair it obviously depends on the situation. If you are churning out large volumes of PDFs then it might make sense to get the efficiency with a PDF-specific language. But HTML -> PDF definitely has its place too and is not as hard to work with as people are claiming.
For common cases, it may be possible to basically decompile the PDF, modify the text, and re-flow the text, and re-compile to bytecode. However, it's very complicated to do in the general case. (Note that in HTML, the browser determines how to best layout the text, but with PDF, the PDF generator makes the layout decisions.)
Also, many PDF renderers will "compress" fonts by lazily building up an embedded font as glyphs are used in the document. These typically will assign "a" to the first glyph used "b" to the second, etc., so if you decompile "This is some text", you'll see "abcd cd defg hgih". Some PDF generators will helpfully annotate the generated text with "backing text" metadata to help screen readers/copying-to-clipboard, but it's far from universal. So, you might need a database of hashes of all of the bytecode functions in a large number of fonts and/or some image-to-text software in order to reliably decompile the PDF.
If you're unable to copy text out of a PDF or you get gibberish when you copy text from the PDF, it's likely because the PDF lacks this "backing text" metadata (and in the gibberish case, likely a compressed embedded font). Some scanners will helpfully perform OCR to add this backing text metadata to the generated PDF.
Source: I did a small amount of work related to PDF analysis in Google's web search indexing pipeline over a decade ago. Most of my work was related to figuring out how JavaScript altered web page text, but I did learn just enough about PDF to be dangerous. At the time, Yahoo was Google's biggest competitor, and tons of their indexed PDFs had preview text that was this compressed font "abcd cd de..." garbage. Yahoo obviously naively decompiled the PDF and just trusted that "a" in the embedded font was a bytecode function that drew the glyph "a".
Maybe out of coincidence all the letters are present. Then you'd have to deal with manually adjusting the spaces and reflow the text. Reflowing the text can be done, but cumbersome. It's akin to fixing a bug in program not by changing the source and recompiling, but by binary patching.
In contrast, it's much easier to delete some letters in the PDF and keep everything else in the same place. In fact I've had obvious PDFs that have a copyright notice on every page. Deleting that can be done with qpdf and just vim (basically deleting the Tj or TJ operators).
This is fascinating. I recommend you read the PDF specification.
"Just run it as a container" is a bit of an industry cop-out for making stuff unnecessarily complex.
Something like this is likely much more efficient than launching a whole browser for each PDF.
I've used PdfSharp in the past to generate product spec sheets in bulk. It worked fine. This one seems more focused on .NET 6 and modernity.
As a web developer this hurt to read. This is a task which is just crying out for a markup language and a stylesheet, not hundreds of lines of declarative C# code.
Even the "complex example" in their documentation looks like the most basic of web pages.
Code is the only sane way.
PDF is an insanely complex spec (I’ve spent more time reading it than most because I need to know bits of it for my job and I just generally find it fascinating). But a lot of devs just need to put some content on the screen to match a template they were given. In my experience, a complete enough markup language allows you to bang out and maintain those templates better than code.
I know it doesn’t suit every need, but it’s just a way of representing the data so it’s closer to the final output than imperative code is. Definitely take your point though about the limitations becoming dealbreakers.
Page media CSS is designed for this although most browsers don't fully support it, PrinceXML is the go to for full paged media support.
IMO they are not fundamentally different, they are both document formats, PDF just a has fixed paged rendering layout baked in while HTML can flow and adjust to rendering target. The main issue is lack of full print CSS support in HTML rendering engines.
Still, to switch back to the previous point, it seems it's more a divergence between using markup or code to design a document. Both have valid usage and benefits depending on your case.
In my case and my apps, I often need to handle complex conditions that fits better imo in procedural code (complex invoices and agreements). On other cases (reports), I prefer to use a markup language.
If your app is a web app this is a no brainer, the users browser could simply do the print or PDF conversion as needed.
I do see a use for more direct libraries in native apps, although if every native client had a browser control with full print CSS support even then it might not be such an issue.
That's arguable, IME (and also a better UX), most would prefer to just get the PDF file which just one click than to deal with additional browser dialogs. No everyone knows how to do print-to-pdf or even know it exists.
Or do you mean browsers expose print-to-pdf functionality as an API?
If you serve a PDF you still need to hit print or use dialog to save, you can use a headless browser server side to serve that if needed.
I do think browser could use better print API's but you not getting around that with server side PDF's unless the server direct prints to on site printers or something.
Right now, we basically emulate this technique w/ HTML->PDF. We build chunks of report HTML with various string interpolation methods and then compose those to obtain our final HTML output.
Raw, declarative HTML is nice if you don't have an undefined # of things to describe with it. When you are looping and projecting domain types into a report, things get a lot trickier.
We build abstractions for a reason. I think we can all agree that templating markup for layouts has been a reasonable success story of the web generation.
I’ve done something similar but in Python and generating Excel documents. I use jinja for templating to create the xml and then parse that and convert to commands that drive the library that creates the final document.
Starting at $3600 for use on a web server.
At the end of the day, it all depends on how you use the technology, doesn't it?
Unfortunately it looks like this uses extension methods for everything, and those are a pain to use in PowerShell. You'll probably want to write the PDF creation bits as a C# cmdlet instead.
Some people don't want, or can't use, external services like this. Security, Privacy, availability, cost... plenty of reasons.
A library works offline, too.
PostScript is eternal
Second, get a grip, .NET is open source for a while, it's getting more popular again, JetBrains entered the stage with a very competetive IDE, it's easier than ever to use .NET via CLI, days of having to use Visual Studio, or any Microsoft product really, are long gone. I don't like Microsoft either but noone's forcing you to use any of their products anymore.