Happy 10th birthday pandoc
groups.google.com
groups.google.com
1. It's a tool built in large part by folks involved in the humanities. John, of course, is a professor of philosophy. (We could stop there, since his contributions so outnumber all the others, but I'll go on...) I'm an assistant prof. of English. It has large contributions from the curator of medieval manuscripts at the British Library, and contributions from a historian as well. And those are just the ones that come to mind -- I'm sure there are others. It's cool to have a tool used by many folks in the humanities, but also in large part driven by their needs.
2. John is a great project lead, and it's an amazingly open project. If you're interested in getting started with haskell, the filters API allows you to really adapt it to your needs. And if you're interested in adding another reader or writer, there's a good chance it can be added.
For myself, I started out using the python docx library, found it didn't handle enough stuff, wrote one myself (covering mainly my conversion needs, not creation) and had it output JSON that pandoc could read. When that worked, I ported it to haskell, and posted it to the mailing list. Within a few weeks, the docx reader was part of the program. It's been a great experience -- both in using haskell, and in playing a significant role in a program I use and love.
(For me, it's also a large part of solving my own productivity problems.)
It's also driven markdown itself to become more feature rich and internally consistent. I'm extremely appreciative of JGM's efforts to product a unified markdown spec[^1].
Huge thanks, John!
[^1]: http://commonmark.org/
Pandoc is a key component in my Markdown-to-Kindle-and-Paperback pipeline. In fact I've soft-published an article about it: https://medium.com/@gabrielgambetta/how-i-wrote-and-publishe...
Amazing piece of software. Not that often a tool dominates an use case so thoroughly that there are virtually no alternatives, because there really is no need. Thank you, John MacFarlane :)
#!/bin/bash
FILENAME=$(cat title.txt).epub
CSS=style.css
IMAGES_DIR=images
# Ensure images get embedded into the document.
for url in $(grep image $CSS | grep url | sed 's/.*url( *\(.*\) *).*/\1/g'); do
EMBED="$EMBED --epub-embed-font=$IMAGES_DIR/$url"
done
# Generate the ePub, including fonts and images.
pandoc \
--smart \
--epub-cover-image=cover/cover.png \
--epub-metadata=metadata.xml \
--epub-stylesheet=$CSS \
--epub-embed-font=fonts/MyUnderwood.ttf \
$EMBED \
-t epub -o $FILENAME output/*.vars
The .vars files are chapters that have their variable names substituted for values, such as: for i in chapter/*.md; do
out=$OUTDIR/$(basename $i);
cat variables.yaml $i > $out;
pandoc $out --template $out > $out.vars
cat $out.vars | sed -e 's/;;/\$/g' |\
pandoc --chapters -t context -o $OUTDIR/$(basename $i .md).tex;
done
A handy trick to avoid duplication (of names, places, or any text really) within the content.Thanks so much to it's creators and contributors!
http://www.redmine.org/projects/redmine/wiki/RedmineTextForm...
With org support in pandoc, i no longer have to start up emacs to generate HTML for my website, so doing that in batches just got a lot easier.
Pandoc is one of those things that I only recently discovered and totally fell in love with: I wrote a bunch of code in literate Haskell for work, so that I could have the logic right, and then ported it to JS to reduce the bus factor of the project. When I discovered pandoc I just ran it on my literate Haskell and the resulting PDF was so beautiful I had to print it out and tack it up on a cubicle wall, like a "this is what they pay me for" type of reminder.
The ability to write and collaborate on memos in markdown, and turn them into fancy PDFs and .docs is seriously world changing.
We've tried other more capable converters but they are an order of magnitude slower than htmldoc. I guess Unicode and CSS support is expensive. htmldoc is also better than anything else we've tried at not putting page breaks in the middle of a table.
Would you say this type of report generation is something than Pandoc would be suited for?
ASCIIDOCTOR has clear and superior table formatting syntax and more importantly it can work with CSV files. This way you don't have to copy paste data into your reports.
Apparently U+3C1 is not set up for use with LaTeX and that seems to be a prerequisite for pdf. It helpfully suggests I try with --latex-engine=xelatex and that fails after a while because xetex.def isn't in the repositories anymore.
MiKTeX comes with an admin tool that has an option to synchronize repositories which supposedly fixes lots of problems, but apparently not this one.
For anybody that reads this and decides to get Pandoc and MiKTeX, I suggest getting the version of MiKTeX that bundles all of the dependencies. In the days of terabyte disks, I'm not sure it makes sense to pull down dependencies a few kilobytes at a time, especially if packages periodically go missing from the servers.
This is a side-by-side comparison of the HTML version and the generated PDF version:
I do a lot of work researching older and historical documents. Some of these are available in only scans, meaning image-heavy PDFs, or rotten OCR transcriptions. Occasionally I get lucky and find a solid ASCII source.
A recent case in point, Dugald Stewart's 1793 "Account of the Life and Writings of Adam Smith", which I'd found in ASCII format.
A couple hours (mostly of finessing the markup) to pandoc, and I had a 72 page PDF, ePub, and HTML versions, the former two posted online.
https://ello.co/dredmorbius/post/-h-b6ek6segi8eamq2g6vq
ASCII Source: http://socserv2.socsci.mcmaster.ca/econ/ugcm/3ll3/smith/duga...
Markdown: http://pastebin.com/LdKXpHdR
PDF: https://drive.google.com/file/d/0B6Q6JFf-mAJPY0tfaXlQWUNwLWc...
ePub: https://drive.google.com/file/d/0B6Q6JFf-mAJPdXgwcm1NbDhQdEE...
(The Google Drive docs should be world-readable.)
That's only one of many docs I've fixed up using Pandoc.
What's most amazing is it's ability to generate beautiful outputs (for printing, viewing, presentations alike). Awesome.
The only issue I had though is that it does not convert certain charset sometimes due to some latex related issues.
to me it's the perfect tool to generate user docs and manuals for software and other cli tools, my best use case is markdown (github flavoured) to man, markdown to HTML and PDF, etc.
pandoc allows you to use different LaTeX templates than the ones it ships with. So, you can take the default templates, add the overrides necessary for French typography and voilà! (hopefully).
In your document? I find that pandoc + french typography = love as it is using LaTeX under the hood and LaTeX have pretty decent rules for French Typography (better than Microsoft Word IMHO).
You are free to use something else or better