Show HN: PDFs from HTML
pdf.math.dev
pdf.math.dev
As an author, my intent is that the content be easily readable to all readers. I don’t see why I should want or get to dictate the layout and aesthetics to my readers.
There are plenty of good reasons why TeX and LaTeX are still the workhorse of scientific publishing in spite of the emphasis on fixed format layouts.
You might not. Does that mean that music should only be distributed in proprietary formats designed to prevent anyone from plugging in an equalizer?
What about the naturally different frequency response curves of different speakers?
What about room acoustics?
Also, digital, free-flow media lose basically all sense of space. PDFs are much better for finding a piece of content again later, because I can remember the location on the page and roughly how many pages into the document.
Though too be fair, reflowing HTML / displaying it without all the added-on cruft also often fails these days. Tools such as Reader Mode make heroic efforts, but also very frequently fail (or are apparently blocked or have sites blacklisted by tools).
Basically, for a bookmark to fully store a position, it would have to store all of the above (and probably more), and it would only be really usable on the same device as long as the underlying content does not change.
And I am almost always on the same device.
Any other scenario doesn't even ensure that the content will continue to exist, let alone have the same structure.
Zooming in/out should not trigger a reflow. Only things like changing the geometry of a page or font settings require a reflow.
It seems you're blaming a document format for a UX problem created by an implementation.
But so much content has images, diagrams, footnotes, sidebars, meaningful indentation, and so forth, that text reflow often mangles or scrambles or relegates to the end of the chapter/document. And not to mention that reading on a phone sometimes freezes the zoom so you can't even zoom into images when necessary.
When I read a published PDF I'm usually getting a presentation that was carefully thought out for legibility, scale, etc. The locations of images, footnotes, sidebars, etc. all make sense.
And I find that reading PDF's on a phone is actually no problem at all, even on my small iPhone SE 2. Just hold your phone in landscape and zoom so the width of the phone is the width of the text column. Generally it works perfectly well.
So as a reader, when an author/publisher takes responsibility for well-organized and legible layout and aesthetics, I appreciate it greatly.
For non fiction I find reflowed epubs sometimes inferior to a pdf perhaps to a degree more aesthetically than in terms of actual usability which is harder to quantify. Below a certain size this has exactly the defects you describe however I find that on a fairly large wide screen in landscape orientation it is quite readable.
For example a screen 6in wide lend themselves well to reading without zooming. This is largish for phones or smallish for tablets.
Regarding dictating layout and aesthetics for practical purposes most of your users aren't actually dictating much of anything beyond screen size, platform, and zoom level. Just because other settings exist doesn't mean most people use them.
For practical purposes there are small screens where text must be heavily reflowed because not much fits on the screen and screens big enough to show a whole document depending on font size. For most things you want to support the first use case if any portion of your users are going to be on phones which is nearly always true.
This doesn't mean that there isn't a case for designing a non mobile version of content especially if its mostly consumed outside a limited and limiting screen or benefits from such.
Forget text for a second, if I want to see fine details in an enormous image I'm going to have to zoom in. I normally adjust font size rather than zooming text but it's nice to have both available.
Zooming is a very fundamental usecases, which is linked to the need to analyse some parts of a document with more detail (i.e., look at a graph, a section of a table, etc).
Moreover, accessibility is important. There are plenty of good reasons why even Apple provides magnifying glass apps integrated into the OS and that have their own system-wide dedicated keyboard shortcuts.
As for paging speed, just try using GoodReader or PDF Expert on an iPad. I can flip through thousand-page manuals and datasheets as quickly as if it were a paper book. And a 12" iPad shows an entire A4 page without the need for zooming and panning.
In my experience, people who dislike reading PDFs have only tried doing so in Acrobat Reader (which is hot garbage, and slow), on a small screen that is wider than it is tall, zoomed in so that only half a page is being shown. That is a sub-par experience indeed.
This is incredibly important, and something that dedicated book readers like Kindles get right, but I've never seen done well in long web pages. Discrete "pages" (that correspond to "screens") make it much easier to find your place as you go to the next page. Note that multipart web pages often have you scroll through each "page" separately, and give you the worst of both worlds. Sure, PDF isn't always best for reading on a computer or phone screen, but infinite scrolling is annoying too.
Printing this page to PDF outputs the whole website, one post per page.
Temporary interruptions yes. But then the location is kept.
Interruptions for a very long time ? I might have to reread the whole chapter anyway...
Pagination or not, both pagination and files provide some degree of spatial sense just as 'loci' and memory palaces.
Edit addendum:
There are theories about which senses are our dominant ones, and how they affect our learning processes. Some may lean towards visual ques in their mental life, others on kinetic or sound. Personlly I experience my mental models as spatial. Even abstract thoughts become situated "somewhere", if not by itself, then by contrast of other things on my mind.
"Everything is a Memory palace."
Needless to say, when I'm deep off in a terminal with something, I don't think I'd describe it as text-based.
I.e. we strive for short functions, we use indentation heavily, it is commonly rendered in fixed-width fonts (this helps with spatial memory/overview too), etc.
Which also seems to align with the article's description of how it works: you're not trying to figure out the underlying structure of what's going on, you're making up a new structure as you go based on the surface-level patterns.
It's the things that didn't need to be PDFs, but inexplicably are, that annoy me. Like data dumps from local governments that could have been machine-readable, or announcements that are distributed in print and emailed as PDFs, rather than lifting the content into the message body.
It's annoying, but if they were produced from a database (as opposed to scans), they're still usually machine-readable by converting the PDF to text, and then running a few regexes as needed to convert to something like CSV, if it's tabular in the first place.
In theory the text could be gibberish because of font subsetting that intentionally scrambles the glyphs, but that's rare and generally only implemented when a publisher is intentionally trying to thwart text extraction and/or font extraction, which I wouldn't expect a local government to either intend or to enable accidentally.
I know of a company that was required to send HR data to a union (time clockings over a period of time). They didn't like it. They just printed a badly-organized spreadsheet to a pdf. There, they sent the data, and it was unusable.
You should thank PDF for giving you any useful electronic copies at all.
If it's scanned-in papers, sticking them loosly in an e-mail or web page would be much more difficult to read through.
If it's text data, then perhaps it was primarily composed to be printed, and PDF allows easy creation of readable electronic copies with minimum of effort from any input. Before PDF you might have gotten nothing at all, because most people don't have readers for various obscure proprietary input formats.
And PDF is far easier than other formats to convert into another format for your own consumption. Do you have a command-line tool which will extract the embedded images out of a Microsoft Word document? Or one that will convert it to plain text, preserving formatting? pdfimages and pdftotext -layout are very widely available.
I think the point is that data dumps in PDF format are not useful at all.
I take objection to your statement that it’s easier to convert. The only reason there are so many tools to do so is because it’s so hard/impossible in the first place.
What exactly do you base that on? Have you written any PDF or postscript utilities?
Images are easily located in rather discrete chunks, and they are conveniently stored in standard formats like JPEG. Preserving the layout of text output takes a bit of work, but otherwise extracting text and images is just about a necessary first early step in writing any PDF viewer. And I do believe even very early PDF viewers allowed arbitrary copy/paste of text.
Conversely, I’ve only ever tried to write a (new) word file once, since it all worked right away.
Links are often hard to pick out. What is a link and what isn't? What happens when I click on something, is it going to stay in the PDF or open a browser or something?
Don't get me started on moving around in PDFs. There are always 2 sets of page numbers, one for the PDF and one of the document. Extremely confusing.
Searching. Ugh, searching a PDF is a nightmare I don't want to even think about right now. Ctrl+F is broken 99% of the time.
Or at least, that's my experience over the last 20 years. Sure, it's gotten better recently, but, not enough to make my mind 'at ease' exactly. Very stressful to open a PDF still, usually.
It's only a tiny subset of all PDFs in circulation, but the LaTeX PDFs I produce using appropriate settings (mainly KOMAScript class) always nail this. The current page number always corresponds to what is printed in the PDF. This can be alphanumerical (e.g. page "a" / 300, where 300 is the total number of all pages) or roman, for the frontmatter. The PDF viewer will then literally show e.g. "Page XII / 300".
So in that sense, it's in the hands of the party producing the PDF to get this right, not an inherent limitation in the standard.
But now, new problems arise. If you're on the printed page XII but your viewer displays "Page 22 / 300", you know where you are in total. "Page XII / 300" is "correcter" but can be anything.
> Searching. Ugh, searching a PDF is a nightmare I don't want to even think about right now. Ctrl+F is broken 99% of the time.
Don't share this experience. It's a the same level as in browsers, where CTRL+F is also quite limited (I'd give a kidney to have regex available everywhere---ripgrep-all gets close on the desktop). The only different thing in PDFs is if hyphenation occurs, which is arguably less common in browsers (simply because of poorer typographical standards/people care more in proper PDFs). Your search term will indeed be invisible to CTRL+F. The only other time it breaks down in PDFs if the PDF is corrupt/poorly produced/bad OCR.
The fact is that PDF can display one thing and have underlying semantic text be something else entirely (frequently used for OCRing: you show actual scanned images of text, and put the invisible OCRed text as searchable ).
It works in the other direction too: you could solve the hyphenation problem in the same way by having PDF include invisible non-hyphenated word in place of the hyphenated one for searching.
Still, PDF is mostly a laying-out format, and while tools have evolved to provide some "meaning" to rendered content, it is never going to be semantic in the sense markup languages can be (i.e. there is no "emphasis", "quote" or "header" command for PDF, instead, it just uses a different font). To put things into perspective, TeX files can be semantic (if a semantic TeX .fmt like LaTeX is used) like HTML/ePub, but PDF is an output format, just like DVI is.
That probably isn't the fault of the PDF, but the PDF reader you're using.
> Searching. Ugh, searching a PDF is a nightmare I don't want to even think about right now. Ctrl+F is broken 99% of the time.
Now this is actually the fault of PDF and how it does positioning of stuff within it - but about 50% of the blame lies with whichever software generated a shit PDF.
Arguably, it's still the fault of the horribly overcomplicated pdf spec -- html manages to do it just fine, with a plain text format, to boot
epub.
>An EPUB file is an archive that contains, in effect, a website. It includes HTML files, images, CSS style sheets, and other assets.
Search works just fine on ereaders.
The flexibility of html is that you can render it however you want, for whatever viewport or feature. If you want pagination, just render it differently.
>But if you read longer documents you want pagination
For long form works, I do actually prefer having them on my kindle, but that's because I don't want to read long-form text by staring at a screen, and I want a lower line width. PDFs tend to be worst case scenario there, because they often render with the assumption that you're trying to read it on standard letter paper.
Also, we've had "pagination" for text files for decades. It's called "less".
Completely agree. Try something that isn't Adobe Reader!
It just feels a lot better to me. It opens faster, it opens at the position I left it at, zooming in and out is fast, scrolling is smoother, and even if I wanted to, I couldn't modify it on accident.
It just feels a lot more reified than something that is responsive or editable.
Not great for mobile, but that's not what I care for at work.
Occasionally these choices can be good, but often I want to resize the text-size to make reading more comfortable which is easy with HTML or an EPUB, but with PDF I can only zoom so much before I must pan to actually read the entire line. Similarly, I think that the creator's font choice is often the wrong choice, it's very common for me to change fonts on an EPUB, but I can't do that for a PDF which is frustrating.
>Haters of PDFs do not understand the human aspects of it
Unfortunately UX is not something HN audience or Tech in general are good at. ( Apart from Apple )
This is especially important if I'm on mobile, because I can create an easily readable PDF in portrait mode that works great with the Android default PDF reader, whereas the obvious choice to open epubs (Google Books) is terrible: it's slow and battery heavy, and requires you to upload the epub so that it can be converted into the native format. (Once the conversion process choked and I somehow ended up with a 1 GB file.)
b r u h
You might be surprised that's not always true.
https://www.pdfscripting.com/public/FreeStuff/PDFSamples/Jav...
PDFs are great when you want to read them on a large enough screen. They are not great on a Kindle or a phone.
I wish there was a way to have several different layouts in one PDF file, so you could have the same content but with different layouts and then your device could select the most appropriate one.
Now 15 years later I have a private stash of websites and wikipedia articles that I can consult by simply pressing command+spacebar (the files are indexed in MacOS search).
To make a PDF file out of a website I currently use Printfriendly.com, but it wasn't always this way.
Back in the days I loved to use Arc90's Readability (a firefox extension). I don't know what happend to that extension though, there are plenty of old HN articles about that Wonderfull plug-in though:
Post from 2010, probably I started using it right after finding this post... https://news.ycombinator.com/item?id=1153343
https://news.ycombinator.com/item?id=3246081
https://news.ycombinator.com/item?id=3243097
My joy knows no bounds !!
I actually ducked for "What happend to arc90.com ?" and found as the 7th item in the list this website: https://ejucovy.github.io/readability/
It still hosts a working version !!!
Okay kids uses these settings and thank me later: * Style: Athelas * size: small * Margin: narrow * Convert hyperlinks to footnotes
Whenever a pages is worthy of saving, press the button for Readability and pres ctrl+P and save to PDF... that's it.
On top of that and the in-browser Markdown renderer Markdeep, I've built a tool for typesetting undergraduate theses: https://github.com/doersino/markdeep-thesis/
And, coincidentally, just a few days ago I've written a blog post about controlling the settings in Chrome's "Print" dialogue with CSS (other browsers don't support many of the relevant features): https://excessivelyadequate.com/posts/print.html
FWIW, my toolchain is currently markdown files in folders. I pre-pend 000-format numbers to file/dir names so they're assorted by ls or tree. Rendering is a bash script that runs pp first, since files include others using !include(), producing a single .md in /tmp. The mighty pandoc of course, to produce a Word doc, which is then the basis for all further rendering. HTML and plaintext are generated from that with pandoc. I was using pandoc to produce PDFs, but switched to calling libreoffice headless to generate PDF from the Word doc, since this seems to match formatting most closely.
Sounds fussy, but it's a few lines of bash, fairly reliable and reasonably rapid.
One outstanding issue is tables. The Word doc always requires reformatting tables for column width, flow, etc. I can't seem to get pandoc to carry the styles effectively from a custom reference.docx file. I'm looking at ways to render tables separately, format by hand, then include them into the main doc later.
It's a typesetting-specific layout engine that supports HTML, CSS, JS, and even styled XML if that's your thing. It independently developed support for all the latest standards... it's not free but it's very good. The inventor of CSS, Håkon Wium Lie, is one of the product's developers.
I used it on an app a while back to add a PDF export feature to a web app... couldn't speak more highly of Prince.
I have a problem at work that requires highly dynamic content to be generated and output to pdf files. Right now, I am using excel template documents. I would love to use open technologies to do the same. Not able to find anything so far that is as flexible and user friendly. The closest alternative is to use OpenOffice Calc documents.
https://www.lesbonscomptes.com/recoll/
As well as this, I have a script which find all .pdf files without a corresponding .txt file, then generates one with pdftotext. Really handy, I can then easily grep -ril or ag -til for contents. One gotcha: the text files have line breaks, meaning matches don't always work.
https://github.com/Mogztter/asciidoctor-web-pdf
The content handoff goes like this: Asciidoc (using defined roles) generates HTML5 (Paged<dot>js polyfills page areas / pagination stuff), CSS styles stuff, and Puppeteer runs a headless Chromium for the pdf render. It's straight from CSS GCPM W3C spec, a flavor of CSS Paged Media, drafts that have been percolating since frickin' 2006 but have never seen browser implementation.
The beauty of this is that you use the same CSS for web and PDF deliverables. Actually, the even better beauty is that you are using two dirt-common technology stacks - CSS and Javascript - instead of XSL or Prawn or some ancient bespoke layout language. With Asciidoctor, for complex print requirements you're going to be forced to either 1) DocBook-XSL via fopub or 2) DocBook-LaTeX via dblatex. The native Prawn-based PDF tool isn't capable of a whole lot of customization without extensions. So web-pdf is a real shot in the arm for those of us that aren't real keen on going back into XSL-FO.
Prawn’s pure ruby implementation of image layout makes it too slow for the graphics heavy technical manuals I write (though I haven’t used asciidoctor-pdf in 12 months.) I ended up drafting documents with 10dpi images just to get it to render quickly enough for layout but even then, adding images turned a 100ms render into a 6000ms render.
Hopefully this problem goes away with a fast web based stack.
I’ve never had much luck with break-after/before:avoid for <h2>. I hope their css or paged.js works for avoiding this common fault.
We're using a sort of hybrid Asciidoc/S1000D approach, Asciidoc markup with S1000D architecture (filenamers, publication modules, data module codes, etc). The art is SVG brought over from CGM, with conditional content (applicability) controlling the images via a new module type we call "illustration control files" that toggle the art based on "applicability" aka asciidoc document attributes and ifdef/ifevals.
PDF is via DocBook-XSL, but it's a scheduled process and not "on-click", which I am positive would break things. I am not even sure how to fire fopub from this company's web architecture; they wouldn't let us post html to a network directory (argh?), so any hopes of doing something more advanced are pretty low. In my off time I am looking at Antora pretty hard, and web-pdf is going to be the default pdf tool for that build platform. One thing I am wondering is how Antora's playbook files are going to relate to the "Ascii1000D" Publication Modules, which overlap a wee bit.
The code is pretty terrible, but you can see an example (my resume) here:
* Data - https://github.com/aviraldg/aviraldg.github.io/blob/master/_...
* HTML - https://github.com/aviraldg/aviraldg.github.io/blob/master/r...
I created a PDF exporter for a manual test tracking app using this -- render to (pretty simple) HTML, pass to the prince executable, and out comes a beautifully typeset PDF.
Prince has its own rendering engine that is purpose-built for PDF rendering. It's actually very good - a lot of professional books and documents have been typeset using Prince.
It's genuinely unbelievable. If the PDF isn't sufficiently structured, it has OCR that seems to "just work".
You can also automate the extraction and integrate it into your pipeline.
The UI is pretty old and ugly-looking, but it is one of the few apps I've used in the last 10 years that made me feel genuine delight.
* From my observation and guess.
OSS Python library to generate PDF reports from HTML, using pagedjs. Uses Jinja templates, supports runtime-generated images, client-side JS, and reports are bundled as a single file.
This could be very interesting for those with Python workflows, thanks for sharing
It's nowhere near as mature as PrinceXML.
Sure, PrinceXML is unmatched. Same goes for the price tag. I know professional CAD software which costs less. Really not doable for smaller offices e.g.
Its astonishing that there is no real, great open source alternative for this. Don't get me wrong, Weasyprint is great, but has _a lot_ of dependencies and is a nightmare to install on Windows. Works decent on WSL, tho.
That's what I like about pandoc + weasyprint. I just have a plain ol' markdown document and I receive a nice PDF in an instant. Just like that, super easy.
Do you have much experience with paged.js? I wonder what the benefits over weasyprint are.
Yes I know what you mean - it doesn’t work well in certain browsers, devices etc. Aimed at desktop users and Chrome in particular (I saw a bug in Safari version). The aim is less to be a readable pdf in-browser, but rather a high quality pdf after exporting to pdf. The in browser print preview is just a nice side effect (but I might actually reuse this for other projects as I like it quite a bit!).
I think the issue with the toc is that it’s dynamically created; so while I was able to use responsive web design for the rest, it didn’t work so well for the toc. I’ll have a look at it though :) I think there may be a way to get it to work
In case of academic research papers typeset with LaTeX, the source file is something you'd likely want to consider the semantic equivalent of HTML. TeX should be able to render the same document with different output constraints ("responsive layout"), but because of the architecture (TeX itself is fully Turing complete), it is pretty slow at re-rendering an entire document.
Part of the allure of a static document format like PDF is that you can, in theory, fetch just page 454 of 6000 page document and render that: with HTML, just like with TeX, you'd have to get and render the entire document to be certain that the layout won't change after you've processed the whole file.
The option to specify the target output size / dimensions at generation time is a reasonable option --- ISO A4 / US Letter, perhaps a target for smaller devices (though a 6"--7"+ tablet should be able to present most reasonably-formatted PDF documents reasonably legibly).
For anything smaller, PDF isn't really well-suited, and your better option is to go with a fluid-layout format such as ePub, .mobi, or, yes, HTML.
Having largely switched to bookreaders (eInk tablets), in large format with ~300 DPI grayscale screens, I strongly prefer fixed-layout formats such as PDF and DJVU to fluid-layout formats, for the spatial/cognitive reasons many others have mentioned in this thread.
My (and others') point is and remains that a spatially-fixed layout does serve a useful purpose for some documents. Including the 140 million or so published books and an even larger count of formatted published articles.
Yes, for short texts, dynamic flow within an HTML webpage is useful. Yes, for very small devices, virtually any format sucks and blows (this is a device problem, not an inherent PDF problem).
I'm not a fan of websites that dump what should be Web-formatted content as PDFs. But I'm also not a fan of the notion that everything should be an HTML document either.
(I've used various online document formats for going on 40 years, from raw ASCII (or EBCDIC) through roff/nroff/troff/groff, HTML, LaTeX, various flavours of Markdown, etc. I've hand typed out several books simply to have a suitable online digital format of them (I hope this serves to indicate my level of obsessiveness, if not sanity, on this topic). I'm a huge fan of Pandoc and its ability to take a standard markup format and produce a wide range of output endpoints (usually: PDF, ePub, HTML, plain ASCII text, though a few others may be included).
I'm also a recent convert to large-format eBook readers. And from that experience I can make two specific observations:
1. The behaviour of HTML and web browsers on an eInk device really sucks. Pagination and not triggering scroll actions with the merest suggestion of a hint of breathing on the surface is hugely underappreciated.
2. PDFs (or equivalent paginated documents, e.g., DJVU) offer an excellent reading experience on such devices.
I'm not a huge fan of the PDF file format, mind, it's far too variable and has too many surprises and vulnerabilities. As a reading medium, however, it's quite good, especially when produced with competent tools.
That said it works in iOS as far as I can test. For some reason page numbers in table of contents not working perfectly in Safari but Chrome works pretty well.
I guess the conclusion is that this is aimed somewhat to desktop Chrome users as a specific tool for pdf generation.
Rest seems to work on Safari - let me know if any other issues and I’ll fix / update accordingly
My go-to for everything CSS Paged Media is [0] which has a nice comparison of supported features at [1]. They recently added Weasyprint, PagedJS and Typeset.sh
Paged.js was a revelation (thanks to HN for telling me about it!). It is based off CSS Print Specifications like PrinceXML (as is my understanding - I’m about 95% sure), and to me it’s even better because it utilizes all the other front end technologies directly from your browser - I think there are some use cases where PrinceXML won’t be able to get the same functionality.
For invoices, I think you should be able to easily switch over. Based on what I can see.
Edit: Never mind. I found the instructions on one of your links: https://github.com/MrRio/jsPDF#use-of-unicode-characters--ut...
https://resumetopdf.com/fonts/OpenSans.json
And here's me embedding it and making it available to jsPDF. NOTE the font.replaceAll to remove white spaces from the font name when embedding. jsPDF has this bug (or maybe expected behavior where it can't handle white spaces in the font name):
var fontBase64 = {}
const doc = new jspdf.jsPDF({
orientation: 'p',
unit: 'mm',
format: currentresume.size,
putOnlyUsedFonts: true,
})
const fontsToEmbed = [currentresume.headingfont, currentresume.bodyfont];
for (font of fontsToEmbed) {
const fontName = font.replaceAll(' ', '')
if (!doesExist(fontBase64[fontName])) {
let response = await fetch(`https://resumetopdf.com/fonts/${fontName}.json`)
let data = await response.json()
fontBase64[fontName] = data
}
Object.keys(fontBase64[fontName]).forEach(style => {
doc.addFileToVFS(`${fontName}-${style}.ttf`, fontBase64[fontName][style])
doc.addFont(`${fontName}-${style}.ttf`, fontName, style)
})
}
Once added to jsPDF, later on you can set the font and style, weight etc: doc.setFontSize(originalFontSizeInPt)
doc.setFont(fontFamily, fontStyle)
Also the JSON with base64 of font families is served over CloudFlare with an infinite cache - this prevents any costs on my end and also speeds up the user experience when they return to the page.Ha! Just checked you account and indeed that's what you do... I'm not sure I like to acknowledge the sense of failed HTML but your pages look good and there's no javascript nonsense in the background.
The repo (https://github.com/ashok-khanna/pdf) contains all the necessary code and is intended for others to reuse in their projects. Some of it isn’t straightforward, despite the guide looking easy - I had to figure out how CSS selectors and counters work for example, how MathJax interacted with Paged.Js.
I think the confusion comes from it being labeled as a “guide”, in fact it’s a full set of code to give the required functionality for high quality PDFs from HTML, using paged.js, the guide is just the self documentation as I figured I might as well use documentation for the sample output. Otherwise, I’d be genuinely curious on what constitutes Show HN vs normal posts?
I think the repo description and the way the output is confusing / unclear - the primary goal is very much meant to be a code base for people to reuse as I’ve noticed for many programmers, the design side can be a bit more elusive.
Separately, would it be possible to add beautiful back to the title - it’s not really about producing PDFs from html as browsers can already do that, and there are many other tools. The main aim is to have the functionality to produce very high quality typeset PDFs from HTML, which until now, I only felt PrinceXML did well and that’s a paid solution. Maybe we could say the title is “High quality PDFs from HTML using Paged.JS”? I know there has been a separate discussion on another thread on the overuse of the word beautiful in describing code - my view is that it has its place when it relates to output / UI.
Thanks for reading, and no issues otherwise (no need to reply).
Apart from this feature everything worked fine.
But yes, it’s a deficiency in the system currently
Reset page number using counter is working: https://codepen.io/julientaq/details/MWammZV
The property needs to be set to the element, and not to the @page.
But we may have missed a bug though, so i’ll be happy to check your code.
#grampstextdoc {
counter-reset: page 7;
}
@page {
margin: 2cm 3cm;
@bottom-center {
content: counter(page);
}
}Would it be possible to add a license so it's possible to know whether others can use this in other projects without rewriting the CSS from scratch?
Also paged.js is the property of their team, together with the script of toc.js whilst MathJax is the property of their team too. Have to figure out how to word it.
But if anyone is reading in the meantime, it is open source, no need to attribute anything back to me for my parts. If you are using the text of the guide, you could mention my name, but don’t sweat it either - it wasn’t particularly involved in terms of writing (the hard part was choosing which parts to write about so it’s not too complex but also not too barebone).
The idea to polyfill the required functionality for CSS Print Rules is pure genius, as somebody who looked at all the alternatives. It’s a great example of thinking outside the box.
Thanks again!!
I may write a book on typesetting with CSS, the quality of what is on the web is not the best, but it seems like a huge time sink at the same time...
This is ideal for reading from a tablet or from a desktop and is somewhat printable too,
Unfortunately most tools for producing PDFs from HTMLs assume you want to divide them into pages, and there is no easily produceable "reading format" as widely adopted as PDF. Those page cuts are so annoying when you read from digital media.
Let me know if that works well?
I've been looking for report generation solution based on frontend technology.
Btw. This is great. Thanks for sharing.
Would you be able to expand on what you mean by “report generation libraries”?
For example, I am building (in Common Lisp, but it’s trivial and can be done in any software) a tool to read content from a database and auto generate the HTML markup for producing pdf reports. This allows me to reuse content across reports and also leverage the full power of databases (text search in particular). As another example, I have many monthly financial metrics - I will store these in a database then use my lisp markup tool to generate the necessary HTML to produce the pdf report (via paged.js).
In addition, one can use headless chrome to automate the full workflow so that the reports are generated directly from your program and not via File > Print in your browser.
Was that what you were thinking of?
You can also add charts via charts.js.
The beauty of paged.js is that you can leverage many of the features of browsers and JavaScript libraries in your report generation.
I wasn’t able to get syntax highlighting for code blocks to work however, need to dig into that a bit more.
Yes! This is great. Thanks for the pointers. I'll look into that.
The actual PDF generation component is an electron application, so it may fit your "frontend technology" requirement.
I generate some HTML for the user, they can edit it with a rich text editor like TinyMCE, and then we export to PDF on the server side. Wkhtmltopdf is pretty barebones on the style/feature side though, so this looks worth investigating.
But maybe that’s because I didn’t fully learn wkhtmltopdf
PrinceXML is very good but is a paid solution. I found paged.js, at least for my purposes, on par with PrinceXML
I was searching for many weeks for something like this, so I really think the word needs to get out there more. It could significantly improve the workflows of many people who are self writing / self publishing as it opens up the power of CSS and HTML (which allows to nicely defined formatting templates and use code to automate content generation) to pdf reports (which I think has its place).
I haven’t used pandoc, but I think a HTML/CSS/Paged.js workflow could challenge it.
At work I’m already converting many processes to it - I have a database of content and then use SQL queries to extract data and then generate beautiful PDFs through paged.js.
It also works well with mathematical typesetting (via MathJax).
Consider adding in the demo automatic anchors on headers so one can quickly copy them for sharing. Currently they can only be obtained from ToC but you need to scroll to it. On anything larger then few pages, this is a must. One problem there is that current automatic id's are generated sequentially and not really user friendly for link sharing.
Would you be able to expand on what you mean (sorry I’m being dense), otherwise I will Google it tomorrow
Thanks for the kind words
If you go there, I can grab it hovering over header (↵ symbol to get permalink). It also shows nicely in the URL by combining words instead of generating sequence.
We solved this by post processing the pdf generated with paged.js + puppeteer with itextsharp (LGPL) to add the bookmarks.
We captured the toc using the paged.js "after" hook and put that into a variable which our backend could then grab from puppeteer.
If you have to render something big, like a 300 page member directory for example, the approach will blow up.
And you can also check https://villachiragan.saintraymond.toulouse.fr/impression to see HD images.
Both went straight to the printshop.
I was thinking about doing it, but it would be a lot of work to do right.
By right, I would want it to be the quality of sublime or emacs / vim :) :)
How about images? How do you handle images; their layout, scaling etc?
If you want floating images (e.g. text on the left, images on the right), it may be a bit more difficult and not perfectly possible. This guide will help: https://www.pagedjs.org/page-floats/
One tricky part is if you want to have text within images and have them the same size as your main text (eg in MS Word where you can have shapes and text boxes). For that, you can probably get close enough with a simple image load, and more precise by using svg graphics, but it may result in a reasonable amount of complexity to make perfect (if at all).
For charts, use charts.js in my opinion.
I would have loved to have something like this for a project years ago.
It’s such a big deficiency in the modern web, we really need Chrome / Safari / etc to implement the W3C standard or something better
Can this be integrated with Wordpress?
Consider the alternatives.
HTML, too much re-rendering and re-formatting.
Word - Oh No. Not in a million lives. Have you seen the atrocity that is the "reading mode" in Word?
Epub / Mobi / Etc - I have never come across good readers.
For what its worth, PDFs are great for reading on larger screens like iPads. I read them on my mobile too, but that's not good for long reads.
The price is pretty reasonable up to about 7-8", and not too high to 10".
I splurged for 13" as I read a lot and often low-quality scans of small, three-column print.
The pixel density of an ebook reader (200-300 DPI) is far higher than even Retina displays. Monocrome/greyscale gives higher resolution as well (the three-elements-per-pel aspect of colour displays means you're always left with about 30% the effective resolution, though subpixel aliasing helps a lot).
Portrait will display a single page well (laptop displays suck for reading text), and for larger devices or larger-print materials, you can often manage a two-page-up display.
As a format, yes, what the fuck is everybody talking about? PDF is a disaster and should be killed off, HTML is great.
Typo FWIW on page 8:
"You want to select the print background option and delect the print headers & footers option."
Just write html + a bit of css (no 3rd-party complexity!) and print it (to pdf).
Done (assuming you won't apply as a print designer)