Twenty Years of Pandoc
pandoc.org
pandoc.org
Beautiful writeup for a wonderful project. In an age of vibe-coding hype it's also so nice to see how things can be extended and snowball in usefulness when things are built correctly, by hand, from basic principles.
> Perhaps, then, in the future, people will no longer have a need for tools like pandoc.
I think we will need wonderful things like pandoc more and more. As mentioned there is a huge ecological and practical difference. Even if LLMs could get infintisamally close to deterministic-level reliability, it's still so many more orders of magnitude better in efficiency, especially with big batch jobs etc.
I feel this influence of choosing a tech stack and its impact on self selected and auto-reenforced culture is most often underestimated.
From my own experience, at a time I was (involuntarily) working in Java, and when .Net was released, from a pure technical point of view it was like a breath of fresh air. Java was suffering from overengineering, archtecture astronauts galore and no sensible UX framework. .Net, the new kid, came in lean and clean with a UX library that 'just worked'.
Problem later was that for all its flaws and being overly 'academic', in teams (the real thing, not the awfull app), you could have indepth discussions about non trivial aspects of SWE topics in the Java world, whereas for all its technical prowess, in .Net land you were mostly dwelling amongst the 2 week CRUD app bootcamp folks. This ofc is a gross oversimplication. You had brilliant engineers and challanged codemonkeys on both sides. But the skew was more than a little biased.
I find this mindset very refreshing in the era of availability often implying noise over signal.
In other ways it made filtering through candidates quite simple: half of the people who applied were really good, the other either language astronauts or folks fascinated with the tooling who didn't genuinely want to move fast to produce commercially-viable software, they wanted to tinker. You just needed to figure out which bucket the person was in.
Language is not the problem. If someone is a senior developer (not just has the title because of years of service), you can teach them Haskell on the job for little cost. Sure it will take them a few years to be an expert in the language, but most problems they need to solve don't need language experts, just someone good enough. And Haskell is a different language, most often you are hiring for a language that is only slightly different from ones they already know.
Pandoc is my go-to tool. Thank you, Sir!
https://gist.github.com/rahimnathwani/210b1f9cb6ce731a304322...
cat email-to-xyz.md | md2clip
clip2md > email-from-xyz.mdtidyhtml () { pandoc -f html-native_divs-native_spans -t markdown-raw_html-raw_attribute | pandoc -f markdown -t html }
which strips out styling, wrapper divs, spans, inline attributes, etc from (for instance) HTML copied from a google or word doc. Just the semantic goodness!
find . -name '*.md' -type f -exec sh -c '
for file do
out="docs/${file#./}"
out="${out%.md}.html"
mkdir -p "$(dirname "$out")"
pandoc --quiet --template template.html "$file" -o "$out"
done
' sh {} +It's very handy to switch when tinkering with template changes because I can convert ~ 1000 markdown files instantaneously on my mid 2010s desktop.
I never heard about djot [0], is anyone using it?
[0]: https://djot.net/
I have, for a while, been toying with the idea of maintaining a fork of (or plugin for?) Overleaf CE that runs pandoc-powered build scripts instead of using latexmk directly. Perhaps this might be something vibe-codeable.
If you compare this with the HTML produced by typst or hevea, this is super useful, as we can then roll out our own styles.
pandoc document.md --to native pandoc -f html -t native "https://news.ycombinator.com/item?id=49156750" | bat -l hs --style=plain --paging=neverAs for Haskell, I guess tree transformations and parsing are the perfect use case for functional programming. I studied in Utrecht in the nineties when Erik Meijer was still teaching there (later went to work at Microsoft Research where he contributed to things like F# and Linq). In short, my compiler course was taught using functional programming. We were toying around with writing our own parser generators to implement a subset of Modula 3 or our own toy languages. Lots of monads and other esoteric abstractions.
I haven't really done much professionally with any of that since except having a really easy time when languages like Kotlin, Javascript, etc. started borrowing liberally from functional programming. These days, if you have a list, calling map or forEach on it with another function is perfectly normal in many languages. Very nice alternative to a for or while loop.
For similar reasons I would rather pick up Swift, Scala, Kotlin or F#, instead of Rust, in the areas I work on.
I use it to create pdfs for my blog posts with that code :-)
python3 "$SCRIPT_DIR/preprocess.py" > blog-post.md && \
/datadisk/Downloads/pandoc-3.9/bin/pandoc \
blog-post.md \
--pdf-engine=typst \
--template="templates/blog.typ" \
-o "out.pdf" 2>/tmp/pdf-pandoc-err.txt;I havent used it much after graduation, but I will always remember if fondly. This was an amazing read, both to learn about its history and as an update on what's been happening lately. Thank you for this great tool and all the work put into it over the past 20 years!
Because, spoiler, it's not.
Constructs like tables and bibliographies challenge natural language markup - resulting in an explosion of different formalisms - but component content system artifacts for transclusion and conditionals shatter any pretense that these file types are "documents" at all. Both of those artifacts must draw formal structure from outside of language, i.e., from their own product / domain. They're parts of a system that make documents, but are not themselves documents or natural language. They are meaningless - or, worse, full of wrong meaning - outside of their runtime environment inside an explicit knowledge domain. Something that newer component content formats like Typst recognize explicitly.
The proof of all this is, as they say, in the pudding. What do people write documents in today? Well, they stick to natural language formats, sometimes they let the document system handle tables in some bespoke way, but conditionals are viewed with justified suspicion. DITA and S1000D projects, and the cursed migrations that lead to them, are sparse and driven almost exclusively by regulatory requirements, or, more often, program offices misreading regulatory requirements[0].
And here we all are in the LLM age, where natural language is being vindicated in ways both awe-inspiring and devastating. While component content systems force an LLM to expand its context window to the entire repository to make sense of any single sentence.
The crap of all this is, this is stuff that computer / information science has known since at least the 1980s. There are papers written about it. But high-complexity component content systems are sellable to non-technical writer groups because they don't see the tripwires in the fundamentals, or they think[2] that their product domain is so structured that the tripwires can be rigged as structure.
[0] No, converting to a pile of S1000D 040As doesn't magically fix your MTAs or your ILS or anything else.
[1] I do realize that proximity to natural language is correlate, not cause. Markdown converts well as a low-power notation whose instances denote values; it reads like natural language because that's what low-power does. The operative variable is whether the artifact denotes a document or a function from configuration to documents. AsciiDoc with `ifdef::[]` and `include::[]` converts every bit as badly as DITA, although without the fundamental nonsense of XSD and Horn's Information Mapping.
[2] Almost always wrongly