Pandoc – A universal document converter
pandoc.org
pandoc.org
I'm using Pandoc to write my PhD thesis at the moment, from Markdown source, using certain filters to "augment" what Markdown can do. Examples:
https://github.com/LaurentRDC/pandoc-plot
https://github.com/lierdakil/pandoc-crossref
More info here: https://pandoc.org/filters.html
I wrote a filter that automatically converts URL citstions in markdown to "real" citations in any style you want - very useful for writing papers without fighting with bibtex and managing bibliographies manually: https://github.com/phiresky/pandoc-url2cite
How did you all structure the commenting on your writing? I find converting to odt/doc before sending, managing all the exported versions with comments etc. becomes quite tedious. But I'm a bit reluctant to force my supervisors to use eg. git+criticmarkup[1]. I would love to hear you experiences!
- In my case, my supervisors mostly had handwritten notes which rendered that point a bit moot. However, when I send the almost-complete draft to a professional copy editor, it was indeed a pain to add the comments. Either handwritten+scanned, as acrobat comments, or word comments, they had to be manually input into the markdown file.
- Everything else worked relatively better. It was a bit tedious to type loooong pandoc commands "pandoc --filter=... etc etc" so I recently coded pandocmk [1] to make my life easier. It's not super well documented (but it's a quite short script, so readable), but the idea is that you type the command line options as metadata at the top.
I've also toyed with using it to process code blocks, as a dead-simple literate programming tool.
You can write in markdown and then convert it to word for your uni.
I read that as "mid-conversion" meaning that he can apply filters while the document is being converted
- Babelmark, a tool to compare how different Markdown parsers interpret the same Markdown input. https://johnmacfarlane.net/babelmark2/
- CommonMark, the first formalized Markdown standard, and now the de-facto Markdown standard. https://commonmark.org/ (He's the first listed member of the team.)
I feel like John is probably the single largest contributor to what Markdown is today, other than perhaps the creator of Markdown. Thank you for your work!
The creator of Markdown hasn't touched it in over a decade and yet decided to throw a temper tantrum because CommonMark dared to initially call itself Standard Markdown.
I'm not sure of the specifics but personally I prefer formats that don't evolve over time. So not changing a spec for over a decade should not be considered pathological but actually commendable, if the nature of spec is complete enough for it's purpose.
I know vanilla Markdown is too limited for some use cases. But that is no reason to "overwrite" it.
CommonMark does not make Markdown less vanilla -- it's not like GFM or one of the other standards that adds support for tables or other features.
From Jeff Atwood (who is one of the CommonMark creators):
> The goal of CommonMark is not to redefine what Markdown is, or change the syntax, but make it parseable and predictable.
https://talk.commonmark.org/t/what-changed-in-commonmark/15
Here's CommonMark's statement on why it exists:
https://spec.commonmark.org/0.29/#why-is-a-spec-needed-
CommonMark's entire purpose is to fix data interoperability.
The problem was there was no _specification_. It was a 'how to use' summary, And each implementation could be (and was) different in subtle edge cases.
The point of CommonMark was to define the specification and stick to it.
I agree with GP, I thought it rather sucky that he objected to using the name.
Various markdowns have extension mechanisms, they always have. That's not what the GP was talking about.
However, the problem is you seem to be generally ignorant of widely known points of knowledge about Markdown.
> not changing a spec for over a decade should not be considered pathological but actually commendable, if the nature of spec is complete enough for it's purpose
100% agree in theory, but Markdown's creator never wrote any spec. when creating it. Initiatives like CommonMark are efforts to specify and unspecified language, not to evolve nor replace any existing spec.
The "nature of the spec is complete enough for its purpose" is the part that's not met, though (at least in many people's minds). The Markdown "spec" (either the description written by Gruber or the `Markdown.pl` file) has ambiguities and inconsistent behavior. My understanding is that there were many requests from the community for this to be clarified, but it never was. So I think a decade of inattention is not commendable in this instance. The CommonMark landing page[0] has some more about this issue.
[1] https://talk.commonmark.org/t/the-logo-and-name-should-proba...
Sure, Gruber didn't allow CommonMark to use the Markdown name, but I feel like that's not a super big deal compared to what he did do. The Markdown ecosystem wouldn't exist if Markdown hadn't been created in the first place! I'm not confident someone would have made something like Markdown if Markdown was never created: AsciiDoc and reStructuredText came out before Markdown but have not been as successful.
Gruber's original Markdown spec lacked formality -- and that's where CommonMark eventually filled the gaps -- but I think that Markdown's focus on user experience over technicality was the key to its success over competing formats and WYSIWYG editors (the real competition). By the time CommonMark came around, Markdown had already seen viral adoption; three of CommonMark's creators are from large companies that were already prominently using Markdown.
tl;dr I think the original Markdown spec and CommonMark are both significant contributions in their own right!
My decidedly working class parents insisted we learn to play piano. Most of my relatives had a piano in the house. I think it was a holdover of the days when home entertainment was self-made.
A friend lived in both the Los Angeles and New Orleans areas. He compared the two as: in LA, the parties of rich people have live music. In New Orleans, the parties of poor people have live music.
And Damgård, mentioned earlier, was born in 1957 Denmark, and plays Danish and Nordic folk music. Postwar Denmark was poor. Perhaps this interview (in Danish) explains why he started? https://www.youtube.com/watch?v=AUF_EkN4Z-g
So, 1) is there a significantly high proportion of people in STEM who are into music than non-STEM? (and not simply some sort of observational bias), and 2) is the major contributing factor to the high proportion because the parents of those people were upper class? (and not some other factor like STEM fields paying enough so people have free time for hobbies.)
Maybe listen to "Juke Box Hero" or "Coat Of Many Colors" for inspiration on how people from modest backgrounds can have the same fulfilling experiences as wealthy people. (sorry - personal soapbox)
It was even part of the quadrivium of medieval education: Arithmetic, Geometry, Astronomy, and MUSIC! A classical education not only taught these subjects, but how they were all inextricably interrelated (or intertwingled, as Ted Nelson famously says...)
FWIW, I strongly dislike "STEM" as a term because it makes no sense to me in an educational or philosophical sense. I see it more as an attempt to lower the cost of hiring engineers and scientists by increasing the supply. For example, compare the funding going into getting more programmers and EEs, vs. marine biologists and paleontologists, even though all of them are STEM.
To clarify "no sense to me", I despise Pirsig's "Zen and the Art of Motorcycle Maintenance" because of its insistence on a clean division between romantic and classical views. I view "STEM"'s treatment of the rest of the liberal arts as being similarly incorrect in its dichotomous classification. Eg, mathematics is important for the humanities too.
But it's clear what _def is talking about by "STEM", and there's no need to suggest we or modern culture are following along with a perversion because the conversation isn't aligned with your personal views.
I was on a modeling project that used scripts to generate hundreds of input parameters, embed them in models, run the models, and produce reports. The inputs and outputs shifted a lot over the course of the project, as we came to understand the domain and implications of the work better. At every update, the changes had to be transferred to a Microsoft Word document that went to the project sponsors.
Pandoc made this easy -- we just added scripts to write out the model inputs as Markdown tables, then embed those tables in a larger writeup, also written in Markdown. Pandoc turned it all into a Word document. Thus, the same toolchain that did the actual work, also drove the final report. I really don't think we could have had confidence all the tabular data was right, had it not been automated through Pandoc.
Am I allowed to distribute GPL programs contained inside a Docker image for on-premise installations? Do I just need to provide proper credit and a link to the source code?
Or is there a commercial license available for Pandoc? (I couldn't find anything.)
UPDATE: I've decided to evaluate pandoc and see if it might be useful for supporting Markdown and Word formats, etc. If it is, then I'll reach out to John McFarlane and ask about a commercial license (or just something in writing), perhaps in exchange for sponsorship on GitHub.
Also what in GPL makes this difficult to use it commercial software? You are even free to sell it after all.
Also using AGPL doesent require to use commercial license, where does that come from?!
Better to just use a GPL compatible distribution method: pandoc has 349 contributors; none of them signed a copyright assignment, so you'd need permission from each and every contributor to use the software in a way not permitted by the GPL.
If you need a freelancer with deep pandoc knowledge, please do reach out. I'm happy to help.
I'm not a pandoc user (so far); and have struggled many times in the past with bugs and lacking features in LibreOffice and LaTeX regarding right-to-left text layout and language-specific issues.
My question: How "trustworthy" is pandoc in handling right-to-left content and side-stepping the minefield of target format issues involving such content? Is this subject getting explicit attention from maintainers?
Core contributors are westerners or Russian (US, UK, Switzerland, Germany, Russia), and we rely heavily on user reports to improve non-LTR scripts and languages. But the goal is to make pandoc work flawlessly for everyone.
/If/ you accept that premise, why do you think Pandoc has been so very successful where perhaps other applications written in haskell have not? The Problem domain (something about writing parsers)? The contributors? The culture? Something else entirely?
Of course if you reject that premise I'd also be interested to hear your thoughts on it in as much detail as you care to provide.
Cheers.
But there still may be some truth to the claim. A simple fact is that smaller mind share -> fewer programs -> less chance for extremely successful projects. From personal experience: it took me three tries and multiple months to get comfortable enough with Haskell to the point that I was able to write my first contribution to pandoc (the org-mode parser), despite having dabbled in functional-style Lisp for years before that. But Haskell, as used by pandoc, isn't difficult. In fact, I often find it easier to use Haskell, thanks to its excellent type system. It's just very different and requires a bit more investment up front, with huge benefits lurking down the road.
Data to support my claim that Haskell is actually easy to use: over 300 people have contributed to pandoc, with over 100 contributing Haskell code. Many of those contributors have never written any Haskell before, but the type system helped them to find their way.
I talked a bit about the whole topic here: https://youtu.be/JpNEIpLtCHs
I don't think that's entirely fair fwiw, it's github ordered by stars, that will turn up things used by programmers for programming in any language. But either way I don't find the refutation convincing.
I'd love it if the premise was no longer fair. That the data really does not support it. I want monad tutorials, there are thousands. That is no exaggeration. I want Haskell applications useful for something that isn't programming a computer - really not much.
I was kind of hoping you'd say something about the parsing problem domain and why that /seems/ to work particularly well with haskell but other domains not quite so much, at least yet, and whether that can be changed or is simply the nature of statically typed, pure functional programming languages (I really hope not).
It's not "successful" let alone "extremely successful" programs so much as "existant" that is the bar that needs clearing first.
Pandoc is great. Haskell works well for those of you hacking on it. I've used it, liked it and thank you for it! It isn't necessary to have an opinion on the topic at all, of course.
Not sure if I'll ever find the time, but I'd like to make the org-parser less useful for Emacs users. The idea is to write an org exporter which produces pandoc's AST JSON format; all Emacs Org settings would be respected that way, the detour through pandoc's parser would no longer be necessary, and remaining parser incompatibilities wouldn't matter for users exporting from Emacs through pandoc. Well, some day...
Pandocs makes it possible/bearable to interact with rest of the world (I’m in the process of moving more things to org).
Being able to export directly to pandoc’s AST Json will probably allow to avoid using other programs to edit content at all! I’ll wait for this day to come; perhaps I’ll even learn enough Elisp to contribute untill then. ;)
Yes, repeatedly, and I'd love to know why you think it matters and what it is indicative of!
Also transforming documents seems like a task well suited to functional languages.
Some command (or commands) that can be wrapped in a script:
> convert2txtViaOCR.sh -i input.pdf -o output.txt
Thanks.
Under US law at least, open source software is commercial: https://dwheeler.com/essays/commercial-floss.html
I would bet many people who use Pandoc have no idea they rely on it. I don't think Jupyter or RStudio make a big fuss about it even though they both use it.
I always ponder whether it’s the most practically useful Haskell tool ever written.
I keep a list of all my skills, experience and education in a YAML file and have a LaTeX template that I clone when creating a new resume. Then it’s just a matter of replacing the template fields with YAML metadata and running Pandoc.
It was also fun to make a groff template to output a file you can open with man too lol
Here is an article where I show how to use Panflute, a library that lets you write filters in Python, and how I wrote a set of filters to automate the tedious parts of writing a complex technical manual:
https://gist.github.com/imarko/ec8f39550662fcd16908b7ec9d100...
Can be changed to use .txt or .md if preferred.
I most often use http://markup.rocks/ for converting HTML to Markdown and for testing that my reStructuredText syntax is correct when contributing to docs.
Pandoc also has a demo web page for trying it out (https://pandoc.org/try/). The demo supports all of Pandoc’s formats and doesn’t require a large JS download, but it silently truncates inputs to 3,000 characters.
Let me know if there's anything you'd like to see that would make it more useful for you!
It's also fantastic for converting my class notes from Markdown with LaTeX equations into beautiful PDFs.
I've been using Pandoc (and make) daily for over 6 years for all sorts of document writing (letter, report, thesis, design doc, performance review, you name it) and solve the occasional "interesting" format conversion problem. Its robust, reliable, fast, and a pleasure to use (and script).
a large thread from 2018: https://news.ycombinator.com/item?id=17855104
Does it handle embedded pictures well?
I don't know how they did it, but somehow they put dependency hell on a completely new level.
Yes i'm sure it's a great tool, but there's a limit how much bloat I can tolerate for a single program.
$ zypper info --requires pandoc
libm.so.6()(64bit)
libpthread.so.0()(64bit)
libm.so.6(GLIBC_2.2.5)(64bit)
libpthread.so.0(GLIBC_2.2.5)(64bit)
libm.so.6(GLIBC_2.29)(64bit)
libdl.so.2()(64bit)
libdl.so.2(GLIBC_2.2.5)(64bit)
libz.so.1()(64bit)
libc.so.6(GLIBC_2.17)(64bit)
ld-linux-x86-64.so.2()(64bit)
ld-linux-x86-64.so.2(GLIBC_2.3)(64bit)
libgmp.so.10()(64bit)
libpthread.so.0(GLIBC_2.3.2)(64bit)
libm.so.6(GLIBC_2.27)(64bit)
librt.so.1()(64bit)
libutil.so.1()(64bit)
libpthread.so.0(GLIBC_2.12)(64bit)
libnuma.so.1()(64bit)
libnuma.so.1(libnuma_1.1)(64bit)
libnuma.so.1(libnuma_1.2)(64bit)
libffi.so.8()(64bit)
libffi.so.8(LIBFFI_BASE_8.0)(64bit)
libffi.so.8(LIBFFI_CLOSURE_8.0)(64bit)
$ rpm -ql pandoc | grep -v '^/usr/share'
/usr/bin/pandoc
$ ll -h /usr/bin/pandoc
-rwxr-xr-x 1 root root 162M Sep 30 13:33 /usr/bin/pandocThe Arch (and some other linux maintainers) have made the decision to package all Haskell libraries as separate OS packages and install those as dependencies when you install, say, pandoc. This model of distribution doesn't really make much sense for distributing Haskell binaries, though.
There's a few reasons for this: 1) since most people don't have many Haskell binaries and the few that people use don't share many libraries, 2) Haskell packages are normally statically linked when building executables.
If linux maintainers would simply build/ship pandoc as a single static executable all these issues disappear.
I'm on Arch, but I was under the impression that Debian did the same thing for Node/Haskell modules. Or am I mistaken?
That's orthogonal to the issue. Even with just one package manager, how packages are created and maintained is a separate task.
So, if the pandoc package has dependencies vendored, dependency hell is avoided regardless of which package manager is used to install it.
If, however, the pandoc package has all dependencies listed as separate packages, dependency hell is created, again regardless of which package manager is used to install it.
So this is a matter of policy, not tooling.
With pandoc and all the haskell dependencies, the only downside is the length of the list of packages when you upgrade. If it was all bundled up as haskell-all I doubt I'd even notice.
For everybody interested in alternative installation methods: all pandoc releases are available as statically compiled binaries for Linux, and via installers on macOS and Windows. Any major package managers ship a more-or-less recent version of pandoc. Compiling is as simple as getting the "stack" tool and running `stack install`.
Anyway, I don't really expect Pandoc to do everything, but when you have both Calibre and Pandoc in your toolbox, it sure feels like you could manage close to anything.
Another example where Caliber compliments Pandoc well is when generating ebooks for sideloading onto kindles. Pandoc can create epubs which Calibre can in turn convert to mobi.
That aside, I find the markdown + additional features (e.g. latex math, inline code eval), mainly as implemented in Rstudio and Rmarkdown, to be the sweet spot of power and convenience of typing and legibility in plain text form. Thanks pandoc!
EDIT: I've also used this workflow for reading RFCs for OAuth and such. It's just basically a small curl piped to say away. Sometimes if I feel like reading an article I'll add a readability like cli tool piped between the curl and say commands. Unix is awesome!
Flawless!
I've used pandoc for pdf generation and ffmpeg for some audio recording/encoding/playback. I can't imagine what I would use imagemagick by itself for though (that I wouldn't use some common image processing application for). What do you use imagemagick to do?
Automate various transformations:
- resize - change orientation or ratio - adjust colors - convert format - do all of the above to generate thumbnails of large photos, in one command
I have heard of others, like git-annex, but not used them myself. I wonder if there are any I just didn't know were.
I also wonder if anything about Haskell makes it particularly suited as the implementation language for Pandoc. It must have a lot of parsers in it, and Haskell is supposed to be good for coding parsers.
There are parser generation libraries and meta-libraries for certain other languages, notably C++. I wonder what Pandoc in C++ would look like. Probably a pretty good parser meta-library could be spun out of such a project.
Apparently I use a few Go programs--Docker, maybe others?--but no Java programs at all, because I delete all the JVMs from my machines without noticeable effect. Likewise, no C# programs, because I have no Mono runtime. Probably no Lisp, Smalltalk, Julia, or OCaml. Some things I run almost certainly are or use Lua, and of course Python, Perl, and even Tcl. I don't know of any in Rust, but it would be hard to tell because of static linking.
I.e. a window manager does not seem to me like a useful practical application, once there is a minimal one already. Fvwm was fine.
Pandoc filters allowed me to transform the AST in useful ways. For example I turned the image tag into HTML figures with captions, used the video tag if the URL was a video, and called ffmpeg to encode the video in another format for browsers that didn't support the other format.
+1 for being written in Haskell, indeed way back when I became interested in Haskell, I think it was noticing that this tool I was using was written in a strange programming language that influenced me to eventually adopted it many side projects and to write a little book on.
FYI, https://orgmode.org/list/87y2jvkeql.fsf@gnu.org is about enhancing Org's syntax documentation. If you have specific needs/ideas that you'd like to share, please don't hesitate.
Been using it with https://github.com/Wandmalfarbe/pandoc-latex-template to generate my documents.
Please comment if there are other nice templates, either for LaTeX or for Doc
However it's not quite done, yet. I'm mostly interested in PDF output, and not having LaTeX was one of the goals, so I use weasyprint for PDF generation. Too bad they are very slow with releases, and I encountered many bugs...
- Style using XSL-FO: Use Pandoc to DocBook, XSLT docbook-xsl stylesheets to convert to XSL-FO, Apache FOP to convert XSL-FO to PDF.
I have installed the latest texlive in home directory.
When I invoke 'sudo apt install pandoc' it requires me to install a massive texlive setup at the system level as part of it.
This is not specific to pandoc but many other packages. I have anaconda3 installed in my home, but image-magick requires a massive numpy/scipy system-level install (ignoring for the moment my bewilderment at why would image-magick require numpy/scipy).
I refuse to put up with this kind of bloated bs.