On why Markdown is not a good, or even a half-decent, markup language
schwarze.bsd.lv
schwarze.bsd.lv
Cool, let’s see what mdoc looks like:
> The following is a well-formed skeleton mdoc file for a utility "progname":
.Dd $Mdocdate$
.Dt PROGNAME section
.Os
.Sh NAME
.Nm progname
.Nd one line about what it does
…
So, utterly incomprehensible. I’ll stick to Markdown, thanks!> The most official reference manual for markdown is the original one written by John Gruber in 2004. It is unmaintained since that time and leaves various ambiguities, such that different parsers tend to parse input somewhat differently in detail.
Its ridiculous that it hasn't been updated, though there have been some attempts.
He's only saying if you want to write man docs use the syntax designed for it... Ok. It's horrible to read and write but it's probably the best thing to use.
OTOH if you want to write simple plain text that humans can read and write with no difficulty at all, but can also make nice richer text if you want, markdown is great.
Unfortunately it doesn’t yet have an HTML export version (I would be okay with a simplified one as well for now, like only do paragraphs, emphasis, etc)
I always need a cheatsheet when using roff or latex because they're too much to remember if you're not using it everyday.
If you want features, use a word processor, or LaTex, or something. Markdown is for documents that need only very basic formatting.
The future of this is probably some AI-based system where you put in raw text and tell the system "make it look good ... no, make it look more like a Medium article...optimize for SEO..."
The original TeX doesn't work this way. That's why people don't collaborate directly in it.
The keyword here is "easily". Of course I could learn the language and go wild. But, hey, I had a thesis to write so I had to make choices: PhD or fancy headers.
Seriously, I still can't find a decent WSYSIWYG latex editor with the UX of the legacy Word equation editor or a graphing calculator. The closest I found was [0].
This "boring" type of use is what makes me so damn excited for AI: reduction of low skill monotony.
IMO Markdown isn't intended to structure documents; it's intended to enhance plain text. If you find yourself constantly tweking at whitespace or special characters to provoke your markdown to behave, or if you are using so many features of md that you have to keep a reference handy, maybe you ought to use something different.
As an aside, I have used Front Matter (essentially a YAML header on markdown) for cases when a bit of structure or metadata is needed alongside a blob of text, and I think it strikes a great balance between writability and structure.
Markdown tables need to die in a fire though; what a wreck.
If you write a plain text document and you format it following markdown principles, you get a simple rich text version of your document for free. Your headings will be headings, your bullet points will be bullet points, and your code blocks will be code blocks.
Now, you can also ‘game’ markdown, writing the minimum possible plaintext to ‘trick’ it into making a pretty document while leaving the plain text ugly.
But in general, the idea of markdown is to start from ‘imagine I have a bunch of plain text files I’ve been formatting consistently for years. How can I make them look nice as rich text?’
I'll note this tension has a long history in CS: https://en.wikipedia.org/wiki/Worse_is_better
And I think it has a lot to do with a very particular notion of good, which is something like "conceptually pleasing to someone who has mastered the space".
If you are solely an English writer then yes. Using md with some languages will quickly turn you into its hater. Certain formatting symbols aren’t easily accessible in some layouts, like backticks for code or ^ for superscript. Some symbols require extra keystrokes like ~ in german. Some layouts will require finger yoga exercises. Basically, any other text markups based on English (including LaTeX) are unfriendly to other languages. Thus an abnormal popularity of md across the apps makes them much less friendly for non-English users. Mobile typing pushes this pain to another level.
For a while, when I originally switched away from my native language QWERTY variant, I would sometimes need æ ø å on the computer and I would Google aelig oslash and aring respectively and find the symbols for those letters that way :p
Just look at some of these other examples the author lauds. I've been writing code for most of my life and I find roff's manpage to be absolutely inscrutable. I don't understand the context in which you'd need to pick between roff and Markdown, they seem like they're solving completely different problems for totally different audiences.
Yet, I had to crank out diagrams recently, and found every single program I knew had become enshittified to the point of not being usable.
I got the job done with asciiflow box-and-arrow diagrams, pasted into a word processing document as monospaced-font paragraphs.
You’re really gonna keep that folder of text files in “very good” LaTeX??
Nits:
Markdown is general purpose but limited. It's a big improvement on plain text files. Thank you Aaron Swartz (RIP) and John Gruber.
I sure don't hand-code doc directly in PDF or PostScript formats, although people used to do the latter before TeX was ubiquitous.
This post feels like a rant by someone who’s just having a bad day. Not sure why it’s even on HN.
...It's a great convention on how to do common stuff in plain text. Plain text. Let me repeat: plain text.
(I'm accidentally doing "a different todo list approach" now due to it, and it is so surprisingly awesome; sorry, busy doing stuff off of my "weird todo list"!)
The clue being right there in the name.
There is only so much that can be done in a day but it is very reassuring to see just at a glance what was actually done. Also, this makes never-ending lists pruning super easy and peaceful - "if this item has been carried over for a month and I haven't touched it then it means it's not needed".
And all I use for this is gedit (which is notepad.exe with syntax highlight if you so desire) and have the .txt files sync to my Nextcloud instance.
(Note: So far I use it only for personal projects mixed with day-to-day chores.)
Markdown is phenomenal. It gives end users simple and easy to remember ways to perform the most common formatting operations but leaves the text in a state that it was be parsed by humans without an actual markdown renderer.
* https://impacts.to/downloads/lowres/impacts.pdf
* https://pdfhost.io/v/4FeAGGasj_SepiSolar_Highlevel_Software_...
* https://dave.autonoma.ca/blog/2020/04/28/typesetting-markdow...
My text editor, KeenWrite[1] (see tutorials[3] for details), converts from Markdown to XHTML. The XHTML is then typeset using the ConTeXt typesetting software to create a PDF. The PDF is stylized using a theme[2].
What Markdown is missing, in particular flexmark-java, is a standard CommonMark extension for cross-references and citations[4].
[1]: https://github.com/DaveJarvis/keenwrite
[2]: https://github.com/DaveJarvis/keenwrite-themes/tree/main/exa...
[3]: https://www.youtube.com/playlist?list=PLB-WIt1cZYLm1MMx2FBG9...
[4]: https://talk.commonmark.org/t/cross-references-and-citations...
* There are six word breaks in your pdf, zero in the latex version. (well, good-nature is a compound word in the original text).
* Notice the "river" [3] of spaces that runs downs from "Enfield"
I have no idea why other fixed layout engines don't adopt the superior text-layout algorithms of latex.
[1] https://dave.autonoma.ca/blog/2020/04/28/typesetting-markdow...
I use ConTeXt. ConTeXt uses the same Knuth-Plass algorithm for line breaks as LaTeX. I've allowed ConTeXt to hyphenate liberally. There is a setting to control hyphenation. See lesshyphenation, morehyphenation:
https://wiki.contextgarden.net/Command/setupalign
> Notice the "river" [3] of spaces that runs downs from "Enfield"
ConTeXt probably has tweaks for this as well. Keep in mind, I typeset Jekyll and Hyde in 2020 and am not a ConTeXt expert; ConTeXt has had countless improvements since then.
> adopt the superior text-layout algorithms of latex
Pretty sure ConTeXt is capable of producing the same quality of document as LaTeX.
* https://wiki.contextgarden.net/Documentation
* https://tex.stackexchange.com/questions/36/differences-betwe...
It does appear that the real heavy lifting in this workflow is still all done by TeX, and that Markdown (with lots of extensions) is more or less a database of blobs to feed the template engine.
I would suggest the question: Is what you have at the end really a Markdown document if it has no context without KeenWrite's extensions or templates? I think the article under discussion is a bit pedantic, but I'm pretty sure the author would say that what you end up with is a KeenWrite document.
Documents written in KeenWrite can be typeset using tools that already exist:
* https://github.com/jfisteus/html2xhtml
I've kept compatibility with pandoc (CommonMark) and knitr a priority; I wrote KeenWrite to replace a shell script that was calling out to those software packages. My Typesetting Markdown blog posts dive into that script and programs:
https://dave.autonoma.ca/blog/2019/05/22/typesetting-markdow...
Further, KeenWrite only supports plain TeX, which means that the typesetting system can be any software that is TeX-compatible: LaTeX, ConTeXt, LuaTeX, ExTeX, KaTeX, etc.
Seriously though I do like the workflow you've put together with KeenWrite. I have a potential application where it might be a great fit to replace some markdown->pdf workflows I have built. In my case, I use a Front Matter header to pass structured data along with the markdown to generate templated HTML which can then be "typeset" with the extensions provided by HTMLDOC or wkhtmltopdf.
Thanks again for posting and replying. Pleased to make your acquaintance.
Likewise!
> I use a Front Matter header to pass structured data
The impetus for developing that shell script (ergo KeenWrite) was using variables in documents. KeenWrite injects structured, interpolated YAML strings back into documents by replacing Moustache syntax (or `r# ...` syntax for R docs). This means that the same structured data can be reused for multiple documents, allowing for single sources of truth and a clean separation of content from presentation. Also, variables can be used in diagrams.
As for Front Matter, data can be passed in from the command line[1]:
-v metadata.yaml \
--metadata="title={{book.title}}" \
--metadata="author={{book.author}}"
Under the hood, KeenWrite places the metadata inside the intermediate XHTML document as <meta> tags inside the <head>. These are parsed by ConTeXt using an XSLT-like syntax that converts them into TeX instructions[2].Do post any questions you may have on the discussions forum[3].
[1]: https://github.com/DaveJarvis/keenwrite/blob/main/docs/cmd.m...
[2]: https://github.com/DaveJarvis/keenwrite-themes/blob/main/xht...
https://github.com/DaveJarvis/keenwrite#other
If that doesn't work for you, please open an issue and describe what you'd like for MacOS.
> Little context sensitivity
... this is close to a disqualifyingly wrong statement.
(I literally work with Markdown on a daily basis as my actual day job. Markdown really does suck.)
The only way in which (La)TeX is not context sensitive is that it's sensitive to the entire global TeX state at all times. There's no notion of grammar in TeX, _at all_. It's hard to trust the rest of the content in there after this.
A single canonical reference and a little restraint and a dl definition would do it for me, regarding man page creation.
So, a "standard" which has been continuously revised since 2014 and doesn't have definition lists except as "use HTML"
Its mostly successful, but not entirely. Wish Gruber would just link to that instead of his old, bugger, out of date spec.
Is it just me, or is a JavaScript interpreter irrelevant to a Markdown parser?
TFA is drawing at straws to dunk on a format they don't like.
#Heading
Here's why allowing arbitrary HTML and JS in your Markdown makes parsing painful
<script>
console.log("</script> is a closing tag for script blocks");
</script>In creating my blog and static site I started using Markdown but couldn't get around the limitations. So I invented a simple file format for tagging text into a tree like structure. It has a formatting style modeled after LaTeX. It has served me well (two websites and two blogs).
I call it an "agnostic markdown format" because no tags are defined ahead of time. The user defines everything and writes python filters to generate the output from the tags. I have lots of python code to parse and generate HTML. It has a macro processor, calculator, git date, and other useful tags.
Here is an article about Tagged{Text}: https://etcutmp.com/taggedtext.html Another website built using this format: https://etcutmp.com/evolve5 Source Code: https://github.com/kjs452/tagged_text Source Code for the second website: https://github.com/kjs452/Evolve5Help
The main python file is: taggedtext.py (the parser). All the other python code are filters to generate or transform the tagged text tree into something else. For example, tt_macros.tt implements the DEFINE{} tag.
Markdown's fantastic for generating pure HTML. However, if I had a print deliverable, or any complicated function like transclusion or conditionals, I'd maybe nudge the user to try out an alternative, Asciidoc or something else.
Why do I say that when you can do so many cool things with Markdown? Well, "so many cool things" is sort of the problem. You build up a unique Mkdwn/Extension/<Insert_Framework_Of_Choice/> stack every time you do a "cool thing", until eventually you're writing docs for what is, in effect, a bespoke parser. Knock out one block of that - say, like, your company's InfoSec department decides that technical writers having NPM is a Very Bad Idea - and your cool thing gets scaled back or broken outright. It's much better for your docs' lifecycle to use just the core of the spec, and not be dependent on a flurry of extensions and dependencies for the business of making docs.
I don't include edge uses, where extensions excel. To take one example, Mr. Rojas' excellent WireViz program (https://github.com/wireviz/WireViz). If that gets knocked out, my document still builds into PDFs.
The price you pay, of course, is a more niche ecosystem; Asciidoc tooling is maybe a tenth or a fiftieth of what's dedicated to the M*kd*n Multiverse. And of course, if you're writing simple docs that go straight to HTML without so much as thinking about a print layout, then Markdown is the easy choice.
The author refers to Markdown as "Abominable" at the end of the page.
I am always skeptical when someone has this level of disdain for any technology. The overall argument that "Markdown is bad" would be stronger if they noted the pros of Markdown, and why the pros are less valuable in aggregate than pros of other languages.
I still use markdown for what it does well, but I feel like virtually every other (lightweight) markup language misses the mark as well.
Last year I started writing some posts that are about this in a very vague way (6 so far at https://t-ravis.com/room/doc/) to help order my own thinking and lay some groundwork for eventually demonstrating where a lot of them fall short.
(Caveat: I'm not suggesting anyone stop using them for any given project that doesn't need to target multiple output formats. They aren't unusable. They just all have problems that make virtually all of them difficult to use as a single source.)
On a side note, the text only glosses over TeXinfo, calling it officially dead based on an e-mail from 9 years ago, while it’s had many new releases since; I’m not sure what to make of that. Would anyone have more insight into that?
How's this?
And for those folks, sure, markdown sucks.
But some of us still view source code as primary and want to "use" the docs in the same context we use when "editing" the docs.
Frankly I'm enough of a curmudgeon that I view the whole conceit of markdown (that it was a concession to the "pretty formatted docs" people that was acceptable to us README dinosaurs) as a step too far. Nothing short of writing docs in proprietary WYSIWYG editors will ever be enough for those folks, and we shouldn't have tried.
Here [1] is an example. The dividers underneath headers should have less top padding, so that it's visually associated with the header and not ambiguously associated with either the body or header (or neither).
[1] https://docs.github.com/assets/cb-55935/mw-1440/images/help/...
While I agree with you about GitHub's dividers being too low, I think it's GitHub's problem. Not related to Markdown.
Eventually settled on using ms instead of OneNote, gets me that local goodness and beautiful PDFs but also allows me to define things and make it a bit more domain-specific whenever needed. Used it for several documents at work, got compliments on how good they looked. Wild for a relatively limited system from the 70's.
This name is too short to Google, and Google thinks I am referring to Microsoft, or even OneNote, when I search for it. Do you have a link to this “ms” tool from the 1970s?
While I was doing that, I found Djot, a language very similar to Markdown (a bit more verbose sometimes) but much better specified (and with some additional features). So I did a Djot library in Prolog too.
It is a simple tool, accessible to many, which is useful for simple tasks.
The chief architect of the system didn't comment on the note, but did as he frequently did with various critiques of mine, indicate having seen it.
Markdown fits this description, and really should be considered on those grounds. And it's suitable for documents of often surprising complexity.
Manpages are a specific document type, which have co-evolved with a specific document processor. I've used that processor, and even, whilst at uni, wrote term papers using it. I've forgotten virtually all that knowledge: roff(1) and kin are obscure and complicated. They're also not especially human-readable, in ways that Markdown, other lightweight markup languages, HTML, LaTeX, and DocBook, say, are.
See also: <https://www.dreamsongs.com/WorseIsBetter.html> (HN discussion: <https://news.ycombinator.com/item?id=8449680>)
________________________________
Notes:
1. Google+, as it happens. Which had its own bastardised and painfully limited markup notation. And no, it didn't kill FB, but it was used extensively by tens of millions.
The submitted URL appears to be the man page for a "mandoc" command.
http://schwarze.bsd.lv/mandoc.7
Also apparently the submission title needs (2017).
If it wasn't for Github, I probably wouldn't be using it at all. There were so many Wiki markup dialects and a Wiki Creole that could serve as a base to make a more featured popular one.
I've seen ones where you put something like a plus at the end of the line, the spaces from Markdown, or something like MediaWiki's format where you just can't do it without falling back to writing in some explicit HTML.
I love Markdown, but I also hate this. The only benefit is that is discourages people from messing with line breaks.
I'll just repeat what I always say in these threads: We're in the second decade of the 21st century, and can send anyone on Earth a custom colored emoji in 8 shades of skin color, but bold, italics, bullets and headings are somehow an impossible technological hurdle. Why? Because all the geeks keep insisting that formatting text using random keyboard symbols is a great idea.
And so anyway, despite nodding along with a few arguments above that point, I have concluded that the author and I are looking for extremely different things out of our text markup languages. I want just about as far from TeX as possible - at that point, I may as well just write README.txt's in plain text that resembles Markdown or GemText, but in a less formal manner. At least it'll look the same level of good/bad/whatever in vim as it will when dumped into an HTML file, or a PDF, or whatever. With TeX, I basically have to render the document for it to be readable.
The fact that the author doesn't have to introduce what Markdown is, could give him a hint about why Markdown is not that bad. It's useable, super simple and stupid. It's so stupid that it's used.
I don't think anyone wants to write the README in their GitHub repositories in LaTeX, or worse, in mdoc. I prefer a markup document I can learn in 3 min than a "perfect" markup language barely readable by humans, requiring people to read an obscure and painful documentation for days/weeks.
Wow. No need to read this rant further.
Use what works for you. That could be Markdown, vanilla HTML, LaTeX, DocBook, (x)roff, whatever. No need to justify your choices.
I mean, it does, its called HTML: The key point of it was to dip into another language if you needed to express complex things, not, add everything to the main language
Well that's not markdown's fault.