Why Markdown Is Not My Favourite Language
wilfred.me.uk
wilfred.me.uk
I had to build some markdown functionality into a project the other day, and so I googled around to see if there has been any update on this. Didn't find anything official, but a few random comments and blog posts here and there seem to indicate that StackOverflow, Github, and Reddit are working together to come up with a spec.
(Apparently John Gruber has disavowed any responsibility for leading the community in this regard, so now it's up to the big users of markdown to figure something out on their own).
He right complained about the heavyweight notation required for a simple link in Markdown, but everything is that bad or worse in most Wiki markup.
Really, the complexity of the markup you need depends on your application and any attempt to make a "one size fits all" markup is going to fail. The needs of a simple commenting system are vastly different from a user-editable CMS that doesn't want to give access to full HTML, but wants robust formatting options.
Can you explain to me why wiki markup is universally reviled?
I find Creole to be great text markup. It feels really natural to use and it expresses the intended meaning whether it's rendered using HTML or displayed in plain text. I really like this statement by Christoph Sauer on the Creole website: http://wikicreole.org/wiki/ChristophSauer
I agree that Creole or any other formatting markup is not going to solve everyone's problems. However, I think that whenever Markdown is a good choice, Creole is a better one.
I was used to most of the conventions from Markdown mostly as-is from writing plain text for a couple of decades before I first saw Markdown.
Whenever I come across Creole-style formats, I cringe - it means I need to think about the markup instead of think about just writing.
I'm not sure how you can say wiki markup is universally reviled without lumping markdown into that statement. You must have a particular flavor of wiki markup in mind.
[1] http://site/
And [title](http://url.goes.here) for in-line URLs.
I write links in plain-text emails as: www.example.com/awesome-video-of-kittens.jsp?referrer=whatever
Note in particular that I don't put http in front, nor do I provide an alternative title.
Anyway it should still understand links without http because if I am typing the link out by hand (say because it is a site I remember) then I would still want it to be a link.
This is baffling for me. You are not new to the Internet, but until only quite recently, every single browser always showed the protocol is the address bar. Only of late did Google Chrome and then e.g. Mobile Safari remove “http://” from the display to save space; and yes, naturally they include this string when copy-pasting as doing otherwise would result in an incompletely pasted URI (protocols are not optional, except by current convention in non-hypertextual media).
One clear advantage of markdown is that its markup is designed to be as unobtrusive as possible to the point that its text/plain rendering is its own source. That's clearly not the case of MediaWiki or even Textile.
Stack Overflow supports an alternative syntax for
this use case:
Our primary landing page is <http://www.example.com>
Again, this is not knowledge that the user can take
with them to other Markdown based sites.
Actually, that is part of the markdown spec. See: http://daringfireball.net/projects/markdown/syntax#autolinkTBH the article reads like an elaborate advertisement for Creole.
Take this for example:
Hello world, there isn't
1. There are 4.
The original markdown treats the second line as a continuation of a line in a single paragraph. Github treats the second line as the start of an ordered list. Why? Because then you can do things like this: # My list
1. foo
...
I've been so torn about how to implement so many of these things. I've had a lot of people ask me about that particular issue, and when I tell them the original markdown behavior requires 2 newlines after a line before starting a list, it was the last thing they expected.The main problem with markdown is not the fact that these counterintuitive things exist, but the fact that they're not consistently implemented. I've spent nearly 2 years trying to reconcile all these differences. See https://github.com/chjj/marked for more crazy markdown nonsense in the test suite.
GitHub's change (for comment fields only) isn't list specific, they just never reflow. They document it here:
https://help.github.com/articles/github-flavored-markdown#ne...
If I did agree with you on this, I could edit my post and change "github" to any of a dozen markdown implementations that supposedly implement markdown accurately and it would be just as true.
But, like I said, GFM could obey the markdown spec with respect to the list and still obey its own spec.
To explain this concept:
Why doesn't
Four
+ Five
+ Six
Simply yield <p>Four<br>+ Five<br>+ Six</p>
And instead yields <p>Four</p><ul><li>Five</li><li>Six</li></ul>
It mistakenly thinks that line is a list - that is the bug, not the newline behavior. You could have both behaviors at the same time without conflict. It could return the first output and still be true to the newline part of the GFM spec as well as the original markdown. Instead, it chooses to do something ridiculous.That said, it's really only made for documentation and is based on the Docbook XML toolchain which makes it an absolute bear to use.
And like all of these text-y markup languages, it has corner cases that don't work well.
> The goal is to make it completely readable in plaintext and not super markup-y.
If that was the goal, then I think it failed. I have used it in the past, and it has more weird rules and syntax than even reST. It is powerful though, I will grant it that.Things that actually do annoy me from time to time:
- In the default implementation, an underscore in the middle of a word causes the rest of the word, and everything until the next underscore, to be italic. Fortunately, most of the actual implementations have fixed this, but that ain't standard. I also don't think Reddit has fixed it yet.
- Using one asterisk OR one underscore for <em> and two asterisks OR two underscores for <strong> feels redundant. There is no need for two different syntaxes to achieve the same effect. I would prefer one underscore for <em> and one asterisk for <strong> (email clients have been doing this for plain-text emails for a long time), but maybe that's just me.
- AFAIK there is no built-in syntax for underline, strikethrough, superscript, and subscript. Maybe we should repurpose some symbols to handle such cases.
On the other hand, there are things the article complains about that I absolutely don't think need to be "fixed":
- HTML filtering/escaping. Seriously? There are plenty of excellent HTML filters in every language. No need to reinvent the wheel. Just pipe the stdout of your Markdown parser to the stdin of your HTML filtering library. In addition, some of us use Markdown to write our own blogs, where we have every right to insert an occasional <script> or <iframe>. You know, like embedding a YouTube video. Different people should have different rights to use one or another HTML feature, and managing such rights should be the job of your app, not your Markdown parser.
- Any sort of Wiki markup, like Creole. The distinguishing feature of a Wiki markup is that it makes it easy for users to cross-reference documents. But in order to cross-reference documents, we need to agree upon some sort of document organization system. But given the very diverse situations in which Markdown is used, is it even feasible to agree upon a single such system? It would seem that the only cross-reference mechanism that is compatible with all the things we use Markdown for is the good old URL, and Markdown can already handle URLs pretty well. Converting URLs to something that references another place within your app should be the job of a pre-processor/post-processor. Don't be afraid to write one.
http://blog.nig.gl/post/48802013022/although-parentheses-are...
[test](http://msdn.microsoft.com/en-us/library/dd904817\(v=office.1...
That'll work fine, just insert a backslash before the character and it'll be treated purely as a character instead of a formatting marker.
Also, he says that StackOverflow supports an "alternative" link style such as <http://www.google.com>, but I'm pretty sure that's a standard. http://daringfireball.net/projects/markdown/syntax#autolink
The html is generally handled like this:
parse the Markdown syntax and render as html, then parse, whitelist, and render the html.
Its primary flaw is translating line breaks into HTML line breaks. Per the article's criticisms, it too lacks a formal spec and supports embedded HTML.
As far as I remember, github's wiki pages used to be in Textile format. I might be mistaken here, but I remember being annoyed all the time by the link syntax.
http://redcloth.org/hobix.com/textile/
For styled web content, I prefer Textile. For beginners who don't need to CSS style content inline, Markdown is better.
Non-techies have problems remembering the format, and with typing it correctly, and many will never realize (or bother to look into) that there is a format, even if you provide helper toolbars and prominent links to help pages.
One of our current sites uses Textile because it started up long before Markdown existed. Textile is much worse that Markdown. For example, a single starting space means <pre>. And the link syntax is awful.
People understand "" and "_" very well. Bullets are fairly intuitive. Everything else requires looking at the format code help sheet.
My suggestion for a better link syntax is this: Anchor in brackets before or* after the link. So this:
[Some link] http://example.com/
and this: http://example.com/ [Some link]
would both become a link with "Some link" as the text. If there are brackets-text both before and after, choose the one after.This is unambiguous for almost all kinds of text, and anyone who is oblivious to the syntax is unlikely to stumble into it by mistake. And it's quite easy to remember, especially since the order is unimportant.
You don't need much more than that, because people who really care about formatting are also capable of learning HTML. I see plenty of non-tech people -- many of them writers and journalists, but also non-professionals -- who have learned the basics of HTML because they are really punctilious about formatting and the aesthetics of text.
I use that to write nearly everything which I can easily convert to latex/html/whatever which is especially nice with different templates.
For Ruby, I like Sanitize (https://github.com/rgrove/sanitize). I'm sure you can find something similar for any major language.
re: inline links - No its not currently part of the makedoc2 spec or implementation. However inline links are present in make-doc-pro (http://www.robertmuench.ch/development/projects/mdp/) which is an implementation which includes some extensions to the makedoc2 spec.
Here's an example with inline url in make-doc-pro:
Click on =url http://news.ycombinator.org Hacker News= to get latest tech infoThe linked article is a little wrong though, the "Stack Overflow" alternative link syntax is a mainline feature.
Any link like http://example.com/ is converted into a link.
http://example.com|Anchor text| is converted into <a href="http://example>Anchor text</a>
I thought that was as intuitive as I could make it.
I can never remember using markdown whether it's [http:/wwww.example.com](anchor text) or (http://example.com)[anchor text]
[Some link] http://example.com/
and this: http://example.com/ [Some link]
would both become a link with "Some link" as the text. If there are brackets-text both before and after, choose the one after.This is unambiguous for almost all kinds of text, and anyone who is oblivious to the syntax is unlikely to stumble into it by mistake.
Already being installed on the machines and minds of millions of users is incredibly valuable for a system. It's fantastically hard to compete with that, which means that many successful systems end up being a local maximum.
You could make a new system that's strictly better if you could get everyone to use it. But before everyone is using it, it's not better enough to get them to switch.
Wiki markup exists to bridge the gap between textareas and fully formatted pages. It was a great idea in 2001-2004 or so. We have decent HTML editors now.
Wiki markup makes it difficult for non-geeks to contribute. Wiki markup is non-WYSIWYG, slowing every edit down with some 'parse' or 'compile' cycle. At some point you will have to allow raw HTML and JavaScript to be embedded within it, and then you have the various nightmares that ensue from context-dependent errors in parsing, output, or even security.
There are a few benefits that markup provides over a decent HTML editor - certain changes can be more semantic, and it's easier to diff. As for semantic changes, I would suggest making up your own tags and attributes, as HTML has always allowed, and expanding them as needed to the desired target formats. As for diffs, this is harder, but some focused attention to the problem could solve it. And diffs have their own usability issues anyway; they are barely readable for computer languages. We probably need a better paradigm for tracking changes to human-readable documents.
> At some point you will have to allow raw HTML and JavaScript to be embedded within it
{{citation needed}}
> There are a few benefits that markup provides over a decent HTML editor - certain changes can be more semantic, and it's easier to diff
You forget about indexing and sending plaintext emails. You'll have to do something about soup of tags surrounding a typical WYSIWYG-generated html.
Besides, you don't need to think about sanitization and malicious inputs: markdown output tends to be mostly harmless. You'll have to disable inline html and custom element classes, though, otherwise it might break your page layout.
i started a decade ago on project gutenberg e-texts, so i've had a lot of time to evaluate my decisions, and to tighten up both the format and my converter, which i have ported to several different languages, most recently javascript, where it runs quite crisply.
i've also coded a number of authoring-tools, including online sites, and offline apps which are cross-platform. some of these authoring-tools have a pedagogical bent, making the z.m.l. format easy to learn, especially since it was specially formulated to be simple to grok quickly. (my converter has no complicated regex black-box magic.) and all of my programs have an instant preview facility.
z.m.l. puts its focus on long-form documents, like books, so output-formats include .epub and .mobi, not just .html. in addition, viewer-apps with high-powered functionalities have been created, again online and cross-platform offline.
for an advance preview, send a tweet to me, @bbirdiman...
-bowerbird
-bowerbird
Just like the other markups, there are nice tools like rst2man, rst2pdf etc, and you can use it for making both slides and reports.
1. Foo
1. Alice
2. Bob
2. BarI agree that markdown's flaws are significant, but the suggested alternatives don't address them either (except turning off or carefully dealing with inline HTML).
http://johnmacfarlane.net/pandoc/demo/example9/pandocs-markd...
I saw a comparison here[1] and good that GitHub supports[2] other markups too. I don't think there's much of a need to move to Markdown other than just for the sake of using something else.
I remember reading Jeef's post The Future of Markdown[3] few months ago and then what ensued was really frustrating [4].
[1]http://www.unexpected-vortices.com/doc-notes/markdown-and-re...
[2]https://github.com/github/markup
[3]http://www.codinghorror.com/blog/2012/10/the-future-of-markd...
I would love it if Sphinx (or if there was a decent sphinx alternative) supported creole or markdown, but it is quite tightly tied to reST.
[1]: sphinx-doc.org
For instance you can use `.. versionadded:: 0.2` to indicate that something was added in a specific version. The builder then render a nice and consistent block that can be styled in whatever way necessary. You can use :ref:`bar` to reference something in Sphinx, :kbd:`alt + k` to indicate a key sequence etc.
We could not have used Markdown for this without making a new dialect of Markdown. Also unlike rst Markdown is very ambiguous and restrictive. There are certain elements you can't use below others. That very, very rarely applies to rst. For instance you can without a problem have code blocks in tables or code lists in tables etc.