19 karma · joined September 25, 2023
[1]: https://developer.mozilla.org/en-US/docs/Web/CSS/@page/size [2]: https://developer.mozilla.org/en-US/docs/Web/CSS/orphans
It seems that these conversion engines are massive pieces of work that require a lot of upkeep, partly because CSS is a living spec but also because of the sheer number of edge cases.
We are already working on SOC2 as this has been a recurring ask, and indeed documents almost always contain PII.
It seems that a better format should exist, but the fact that PDF is the de-facto for portable documents make it unlikely things can change overnight.
This may come at a later stage once we have built our own rendering engine though
This may not mean success, it means that game is not over in the documents field :)
We also hope to keep the focus on the PDF generation part rather than expanding super-horizontal style to provide all imaginable PDF tools at the expense that none is really good.
Your second point is very interesting, seems like some kind of .assert('text').isVisible() API. We may want to dig into that further!
CSS actually implements the break-before property to control this https://developer.mozilla.org/en-US/docs/Web/CSS/break-befor... which is also supported by the Print to PDF dialog in modern browsers.
CSS actually implements the break-before property to control this https://developer.mozilla.org/en-US/docs/Web/CSS/break-befor... which is also supported by the Print to PDF dialog in modern browsers.
The way we look at it is PDFs allows embedding of other files and metadata. It is easy to provide a platform where we can enrich PDFs to display different contents than the one in the PDF itself. If this gets interesting enough, we can then phase out the PDF in the first place. But this is a long way ahead.
- We do not force PDF/* profiles down to the user, but it seems that for most of them PDF/UA-1 would be a sensible default. We can extract most of the tags from the HTML semantics by themselves which makes it much easier.
- We target the PDF 1.7 spec. Color profiles can be changed and you can use a custom .icc profile, with the corresponding embedding restrictions based on the document format. MediaBox is supported through the @page size property. Bleed, trim and marks can be added using vendor specific css properties. We don't support ArtBox yet but this is something we can look into! So far none of our customers really wanted to take this out to a real print shop, but we would be glad to help people go down this route :)
In the end, what was the main decisive factor is the support for the PrintCSS and PagedMedia specifications, which have been completely discarded by major vendors and only implemented by specific engines.
Where things differ is that we don't actually use a browser under the hood. This allows a much better control over typesetting and layout - and you can do it on the server. We have also more controls over the outputted PDF and the ability to use more advanced features such as form fields or embedding other files and metadata in the PDF.
We like LaTeX, but even for advanced users laying things out can be a difficult thing. Given that documents are a frontend, we wanted to bring the same tools frontend developers already use.
Edits and corrections on generated PDFs is not provided as the PDFs are signed as-is, however you can attach the metadata to the PDF and rerender with the modifications.
You can have a look at our (WIP) set of templates at https://react.onedoclabs.com/ui/templates where the images are automatically built from the PDFs themselves.
We are trying things out to see how we can make a live preview for development purposes but the challenges of pagination are quite hard to solve in an elegant way at the moment. We are experimenting with Taffy to see how it could fit our use case but this is still a very early tentative.
You are absolutely correct and most existing tools do leverage a browser (sometimes headless) to convert HTML to PDF. However, browser's CSS print specification implementation is severely lacking and layout options are poor to say the least.
There is a second option using libraries that abstract part of the layout process such as react-pdf but although it uses the JSX syntax, you can't port existing HTML components easily.
Onedoc is able to take HTML + CSS as well, quite similarly to what you can do with Resend. React is mostly an abstraction layer that allows you to take advantage of all the existing SSG toolset (e.g. charts, existing frontend components, ...) without having to write things from scratch. The process is thus indeed similar to Astro if you omit any client: directives.
We have put up a small comparison at onedoclabs.com/why-onedoc to show a bit better what capabilities this opens.
Hope this clarifies things!
We experienced the problem first hand creating hundreds of personalised marketing materials and working with various clients: PDF is broken. But it is still needed and we wanted a better way to handle it.
We are launching Onedoc after pivoting away from the AI space, and would be glad to hear what you think is right or wrong with this approach!