A Formal Spec for GitHub Flavored Markdown
githubengineering.com
githubengineering.com
CommonMark contains this little sentence to work around its specified behavior, which is left untouched in the GFM spec:
> A renderer may also provide an option to render soft line breaks as hard line breaks.
I'd say whether or not it does this is a rather important thing to mention. When I write Markdown documents for GitHub I have to change my editor settings, only because of this.
[1]: https://guides.github.com/features/mastering-markdown/
[2]: https://help.github.com/articles/basic-writing-and-formattin...
> --hardbreaks Treat newlines as hard line breaks
We turn this option on when rendering issues, issue comments and so on, but leave it off when rendering blobs (such as README.md). Both are GFM, one just uses this option to make it more conducive to communication.
Do you mind to make sure this is reflected in docs?
> There is a fundamental difference between these two kinds of content: the user comments are stored in our databases, which means their Markdown syntax can be normalized (e.g. by adding or removing whitespace, fixing the indentation, or inserting missing Markdown specifiers until they render properly). The Markdown documents stored in Git repositories, however, cannot be touched at all, as their contents are hashed as part of Git’s storage model.
Conversely, they don't expect the same from Markdown files stored in their repository (e.g. I put each sentence in a paragraph on its own line for my blog, for easier diffing and editing). Additionally, we couldn't normalise these documents even if we wanted (to prevent everything breaking by being over-vertically spaced). Hence not changing the option being not used in this case!
Well, actually it does say that "Hard line breaks are for separating inline content within a block."
https://github.github.com/gfm/#hard-line-break
So two spaces at the end of the line for a <br>, and an empty line for <p>.
https://talk.commonmark.org/t/issues-we-must-resolve-before-...
[2]: https://blog.codinghorror.com/standard-flavored-markdown/
I'd hate to see them pushing a different spec around, that would solve nothing
The first two. [The GFM spec][1] is literally just CommonMark with a few extra extensions added. They even highlighted the new sections green in the spec to make it clear where the GFM spec differs from CommonMark. Everything else is word-for-word identical.
Also, CommonMark failed? News to me. Last I heard it was still under active development, years after the drama with Gruber.
I do, when he has abandoned his project's raggedy implementation yet defends the trademark viciously.
Given this news from GitHub, it's very exciting.
I suppose since one of their goals was to get Github and StackExchange 'out of the markdown business' but neither use CommonMark, and further, Github now has put work into creating their own spec, they failed in that aspect.
A big aspect of the post is talking about how GFM is now a set of extensions to CommonMark. This is a huge win, not a failure.
FWIW, StackExchange [uses CommonMark][1] for their new StackOverflow Documentation site and has [been planning][2] to migrate Q&A to CommonMark for some time now.
[1]: https://meta.stackexchange.com/questions/125148/implement-st...
I think its quite difficult to do though.
Then it was done in OMeta [1].
Previous discussion: A full TCP/IP stack in under 200 LoC (and the power of DSLs) [2].
There's also a PNG parser (but it is not parsing any documentation) in 20 lines of OMeta [3].
[1] http://www.moserware.com/2008/04/towards-moores-law-software...
What you do with the parsed tree is up you though.
;)
For instance it is the reason why there's no reimplementation of TeX.
Nobody stops you from translating the Ruby or whatever into your favorite language.
Fundamentally, markdown was specified as some pattern matching and English description. The original specification was not done thinking of productions and grammar rules, and you typically don't get there by accident.
See http://roopc.net/posts/2014/markdown-cfg/ for a detailed exposition of how the '*' character in markdown is sufficient to ruin any chance of a CFG.
I have sometimes pondered how to make a markdown-like language with a simple production based grammar. I have not succeeded and would appreciate any pointers. The criteria being that is has to have something like the minimal intrusion into the prose of markdown.
https://talk.commonmark.org/t/commonmark-formal-grammar/46
See in particular these comments by maradydd:
https://talk.commonmark.org/t/commonmark-formal-grammar/46/1...
https://talk.commonmark.org/t/commonmark-formal-grammar/46/2...
https://talk.commonmark.org/t/commonmark-formal-grammar/46/2...
I'm not sure they ever came to a conclusive answer (at least on this thread).
Edit: Here is JGM himself saying he doesn't know: https://talk.commonmark.org/t/commonmark-formal-grammar/46/3...
Overall this a step in the right direction but the whole saga is a perfect microcosm of our understructure cranking out pooly-understood stuff which comes back to bite us and cannot be tamed.
Why this? This is not a working blacklist to prevent XSS (e.g. onload="...")
A document that contains these tags will not be parsed properly by an HTML5 compliant parser; the parser will "swallow" other chunks of Markdown content that come after the tags. Hence, we disable the tags altogether.
This is an UX feature, not a security feature. XSS prevention, and a plethora of other security checks, are performed by our user content stack -- but this functionality is shared for all markup languages in GitHub (MD, RST, ASCIIDOC, ...), so it's not discussed in this spec.
Updated my original post to be more clear.
I do think rST is waaay more expressive, but I also recognize that in many of the instances one would want to use markup in a chat or PR situation, the expressiveness likely wouldn't be well received if the trade-off is verbosity.
This is something in life that I file away with competing regex standards: my brain just has to switch languages based on the app in which I'm typing (between markdown, pseudo-markdown (ahem, Slack), org-mode, rST, etc).
I am big of markdown in case I didn't make that clear. I love it
Hopefully now that Github has standardised their own flavour of it (and quite a nice flavour too), more people will start to use it.
Of course there is the obligatory XKCD: https://xkcd.com/927/
At the lowest level, you have commonmark. Then, you have extensions at the top, such as GFM.
If Pandoc/Github/Reddit/SO/kramdown switch, that accounts for almost all front- and back-end cases that I care about.
And given that the first four were actively involved with commonmark, I would take as given that they will support commonmark or a superset of it.
+new text+
could be used...