Using Web Technologies to Print a Book
richardmavis.info
richardmavis.info
GitBook is targeted at people creating technical documentation in Markdown. It has the advantage of git integration so it’s possible to branch & merge your way through a book project. Their GitHub account still hosts the legacy open source editor and cli tools, although the project has moved on to focus on the commercial web-based editor. https://github.com/GitbookIO?tab=repositories
PressBooks is targeted at authors who want to self-publish and has a nice set of book templates (styled by book genre) and a workflow for selling print books on Amazon.
Scrivener (mentioned by dangoor in another comment) is a writing tool with advanced features for different manuscript formats (screen plays, novels, etc) and has cool features for organizing notes, ideas, and references.
system("multimarkdown -s ../#{part}/story.md | sed -E 's/_([^_]+)_/<em>\\1<\\/em>/g' | sed -E 's/<h1 .+<\\/h1>//g' | sed 's/<p>%<\\/p>/<p class=\"section-break\">%<\\/p>/g' > output-#{part}.html")
IMO: s_<p>%</p>_<p class="section-break">%</p>_g
is easier to read than: s/<p>%<\\/p>/<p class=\"section-break\">%<\\/p>/g
Any reason to keep the %-sign?Those sed commands could probably be reduced to one invocation of sed(1) instead of three.
system("cat #{htmls} | #{wkhtmltopdf_cmd} --footer-html footer.html - body.pdf")
No need to use cat.Granted this is Hacker News and people here _like_ writing code (myself included), but sometimes you just want to use a tool to make a thing :)
There's an open source clone called Manuskript though.
No CSS implementation "for the web" does.
Implementing CSS justification for printing as Knuth-Plass and implementing so in a web-compatible manner is somewhat very different task. I think the algorithm's interaction with CSS floats is unclear. As always, it's in the Firefox's bug tracker, no, it's not because browser vendors don't want to or are lazy. https://bugzilla.mozilla.org/show_bug.cgi?id=630181
As always, Unicode is a problem.
1: https://xmlgraphics.apache.org/fop/
Simon Pepping did some work done on Knuth-Plass in FO, but it's old and probably not quite as relevant now:
- https://web.archive.org/web/20070114211331/http://www.leverk...
- https://web.archive.org/web/20070128145517/http://www.leverk...
Markdown → Pandoc → HTML → weasyprint → PDF works great, paged media support in weasyprint is good enough and much, much better than in wkhhtmltopdf.
That is, you probably want to design PDF output. Some users prefer to do any design work whatsoever using web technology. Therefore, PDF needs to be generated from HTML.
I'd probably use Asciidoctor which gives more flexibility on the markup side already, but if this process works for them, why not use it.
In fact, he's not even aiming to design a professional-looking PDF document, just a throwaway printout for proofreaders who don't know how to parse Markdown.
This rendered HTML could be converted to PDF and printed or the self-rendered HTML itself could be printed directly.
It also looks like OP's book was split into several Markdown files, one for each chapter. So he would have needed some sort of build script anyway if he wanted to use texme on the combined document. He would also have needed more than a single line of code in the header, since he wanted some custom styling for blockquotes and code snippets.
True.
> and is unwilling to learn how to design anything if it is not web technology.
False and does not follow.
pandoc makes use of Latex, which is extremely bulky and brings with it a lot of complexity and idiosyncrasies. I was not able to get anything as nice as a GitHub rendering out of that.
The only alternative I know is using web technologies (chrome headless) to produce a PDF, which is what the author does.
I've written several booklets that I sell on Amazon and my website. If you're interested in writing but not ready to commit to a full novel, try publishing a small zine or two. It's a lot of fun.
My latest is a series that teaches JavaScript by creating computer art. Readers learn by copying.
PS. sorry to everyone else for the off-topic!
Oh god, I in my time tried to tweak LaTeX so I could reach the same aesthetics as using InDesign. While doable (I'm sure) - it's not really worth an effort. LaTeX and it's ways of working are very specific to that one context. it's the Torx screwdriver for the torx screws of technical and scientific publishing. But lot of layout stuff needs a philips head, a nail and a hammer and so on. While a torx screwdriver surely can be used to pound nails, I would not suggest it as an efficient tool.
Don't feed the LaTeX fetish. Some people like to do everything with it, just like people like doing lots of odd things that bring particular aesthetic joy to them, that would be completely impractical or intolerable to others.
LaTeX in a non-technical or non-scientific context is an eccentric quirk. I love eccentric quirks and people who have them! I have many myself! But I would not push my quirks to other people in any setting.
... or a dozen, each with different features, like tables.
Phantomjs (and its ilk) are based on browser engines and just don't support this. Also I would love to be able to change layout or content based on where particular elements turn up.
Basically every change to the CSS or DOM requires a reflow (or clever optimizations to avoid that).
The basic idea is to add (placeholder name) stopUpdating/resumeUpdating to window, which can be polyfilled as no-op. The semantics is that CSSOM view methods are allowed to return the value when update was not stopped, or any later value. That is, current web standard forces you to do things "live". New methods give option to do things in batch.
In any case, html/css for paged media should be mostly separate from website code. "Printing out" web pages works in many cases, but it's crappy.
If you were designing something from scratch, such approaches would be worthwhile considering, but I think that boat has sailed, and the architecture would fight against you.
Then again, I believe it was generally accepted that web browsers were stuck with UTF-16, until Simon introduced WTF-8 for Servo.
Yes, modifications may cause a reflow. So? That just means that it’s slow. That’s not a problem.
That’s how you implement such things. The initial implementation throws away all layout information as soon as you modify any CSSOM property, and recalculates it. You release that to people saying “it can now do this, but it’ll be extremely slow; let us know what sorts of things you do with it and when you find particularly awful performance cases, and then we’ll look into speeding it up”. Then, as people try using it, you determine where it’s worth putting effort into speeding it up. This is exactly what Michael Day of Prince said they’d do if/when they implemented CSSOM, when I asked him about whether it might come, several years ago. This is an entirely reasonable approach.
In general I think Tex/LaTex is the way to go in terms of reporting and generation of pdf. The biggest problem with Tex is that it is so different from HTML, and it gets progressively more different and difficult if you have specific layout or style requirements.
What I wish for is a replacement for LaTeX, based mostly on web standards, extensible in javascript... Unfortunately I don't have the resources to do that.
I would suggest not to use element-heavy background designs for as simple a case as this website's. I have a low-grade, student-level laptop, and it makes scrolling noticably lag, by a few FPS.
Other than the unsolicited advise, I really like the workflow. I did something similar for almost all of my papers in college.
• It completely mangled the kerning, like it was ignoring the font’s kerning and then making it even worse by only placing characters to 1pt precision (at 600dpi, one dot is 0.12pt). (To clarify: I never actually measured it; this is just my rough guess as to what may have caused it.)
• It was somewhere between agonisingly difficult and impossible to actually get precise sizes; print A4, for example, with your body carefully set up so the widths add to the right amount, and the appropriate “don’t zoom” command line argument, and it’d still mess it up (and subtle content changes could make it better or worse). A container of `width: 15cm` could end up 15cm wide if you were extremely fortunate, but was more likely to be 14cm, or 17cm, or something like that. And it might vary from page to page.
• Pages didn’t really exist, in layout terms, so that any sort of finesse of where things should appear was just impossible.
• Probably worse, you could end up with the descenders of the bottom line of text on a page at the top of the next page. I have a vague feeling I hit a situation where a line could even be split in half, rather than just the descenders, but that may have been printing from Chrome or Firefox at a similar time.
• Its header/footer stuff was mildly limiting and fairly annoying to get working properly (and made document sizing even more troublesome, too).
It was also very crummy for producing a PDF for screen use, as regards things like links and tables of contents and other annotations.
Had it been just one or two of these things, I would probably have filed bugs; but it didn’t look as though there was much interest in actually fixing things, and it was so very broken for any sort of precise, serious work, that I just gave up.
Have things improved for wkhtmltopdf since then? I’d be interested to hear.
I found the state of the art for web-to-PDF conversion to be Prince (https://www.princexml.com/) by an enormous margin, with it producing absolutely superb results. Nonetheless, it does have some limitations; most notably, in my opinion, CSSOM, so that the JavaScript doesn’t interact with the layout at all. There were various other CSS and JavaScript niggles that I hit too, but they’ve steadily been fixed over time too. Bear in mind that Prince is made by a small team and is the entire web engine.
I would really like a vector graphics pipeline for Servo: https://github.com/servo/servo/issues/3788
[0] https://gist.github.com/shaunlebron/746476e6e7a4d698b373
>Why does markdown do this? If I want two lines to run consecutively I won’t introduce a newline.
hello, wor
ld!
(albeit this is assuming the terminal is only 10 characters wide but I hope you get the point)So it used to be common for developers and sysadmins to manually wrap to 80 characters (a habit I still regularly catch myself doing even now). Obviously with GUI readers and variable-width characters, you wouldn't want that 80 character manual wrap honoured.
Ignoring those line breaks is desirable when reflowing the text into e.g. HTML output