Djot: A light markup language by the creator of Pandoc and CommonMark
github.com
github.com
I like the ideas behind this but I do not think this will take off as much as CommonMark has. My main reasons for thinking this:
1. Users don't really care about edge cases. CommonMark (cmark-gfm really) is "good enough". If something looks odd people go "oh, strange" and make a 3 second edit and forget about it.
2. There's already multiple parsers for every language for markdown. Here there is a single Lua implementation.
3. GitHub is the king.
If it cannot render in GitHub it's not real. If it's harder for me to write in the 99% case then it's not worth it. If it's harder for me to implement within my application then I'm too lazy to switch.
Nice things in here:
1. Less ambiguous parsing. For the parser infrastructure I support at $JOB this makes supporting questions from users a HUGE pain in the ass.
2. AST-first design. Seems like everything appears in the AST (even references!) which is a huge win over cmark-gfm. There's also source positions included which is a MASSIVE win for tooling.
3. Custom Attributes implementation. This is really nice for extensibility. It is less readable but having it as an option is nice.
> If it cannot render in GitHub it's not real.
This is 100% accurate. If I limit myself to CommonMark, with maybe a few extras tossed in gently, I can be pretty sure that my meaning will be correctly rendered.
Anything else, no matter how nice the syntax is, or how powerful the expressions are, won't render for the people I want it to renfer for.
3) Github has renderers for multiple formats if this language ever took off, not unreasonable for them to add one to the pile.
I would agree if it had tables
Instead, this project appeals to me as someone who’s already “bought in” to the pandoc ecosystem. Pandoc makes it really easy to write filters[1] and to take the same source file to generate web pages[2], Reveal.js presentations, Beamer presentations, and long-form PDFs[3]. As someone who writes most things in Markdown compiled via pandoc, I see the cracks in the edges all too often, but I’m too stubborn to give up markdown or any of the tools I’ve built up around pandoc and pandoc markdown. I could absolutely see there come a day where I find some last straw where I can’t get done with pandoc Markdown what I need to get done, and djot seems like it would at least be a contender. I’m sure there are many pundits here would chime in and say “just use Asciidoc,” but every time I look at a syntax quick reference, I get about halfway down the page before thinking “nah, this looks too foreign, I don’t want something that diverges this far from Markdown.”
Djot deviates in annoying ways from Markdown, but not as many and so it’d be an easier pill to swallow for the narrow audience of people like me who want something mostly similar to Markdown that works well with pandoc and avoids the most common syntactic oddities of Markdown.
[1] https://github.com/jez/pandoc-sidenote
I'm curious which of the deviations you find annoying. Djot seems to be pretty sensible to me.
Especially the blank lines around lists restriction. The syntax goes to explain how “tight” lists do not have their elements wrapped in paragraphs, while “loose” lists do. This admits that lists in the rendered output will want to use whitespace to achieve a particular layout. But for the case of a Markdown list like this, which will be rendered with no whitespace between items:
- top
- inner
- top
- inner
- top
- inner
In djot this must become: - top
- inner
- top
- inner
- top
- inner
which has way more whitespace in source form than the rendered output will have. This makes it harder to judge the visual weight of a piece of text by simply glancing at the source, and would require spending more time looking at the rendered output while drafting.And to go further, a loose list would require even more whitespace:
- top
- inner
- top
- inner
- top
- inner
If there were a way to relax this constraint about new lines for list sub elements, I might even consider switching to this for some documents today, but absent that it’ll have to be once I bang into a few more markdown ambiguity problems.While I also like GitHub, I make very little use of its web interface. I mostly use gh to interact with the service. But I use markdown to generate document fragments on a daily basis. Claiming a technology is going to fail because the one big player in town has not adopted it yet is... not very open, to say the least.
There's also on-prem Bitbucket, which supports an even more limited form of Markdown. One could always use CI to render the documentation into Confluence (because it's a safe bet someone who uses Bitbucket uses Confluence as well) using Pandoc, but it's still a downgrade in usability: you can't just review the changes to the documentation done in branch A in rendered form.
But yes, GitHub is king. I don't like mermaid.js, but when I do, I do my diagrams in mermaid.js because GH renders them.
Markdown started with a single implementation AFAIK. Or was there a galaxy-wide coordinated software release for Markdown parsers and I missed it?
> GitHub is the king. f it cannot render in GitHub it's not real.
This can't be more farther than truth for me. I don't care if it's rendering in GitHub or not. If I can use it in my docs, write it fast, put on a webpage in a nice manner, I'm done.
Markdown is not popular because it's embraced by GitHub. It's the exact opposite. If it can capture some mind share, tools and websites rendering this will proliferate.
Personally, I care about the edge cases. I also don't care about multiple parsers because I want to write my own. (I currently depend on Sphinx, Breathe, and MyST, which are heavy dependencies.) And I don't store my code on GitHub because of Copilot and because of [1].
I decided to go all-in on djot for several reasons:
* I can write my own parser.
* djot can target any format, which means I can use the same docs to generate manpages, a docs website like [2], and perhaps PDF's if my docs include something like the Rust Book.
* djot's extension story [3] is the best of any format, and I need extensions that don't exist in any format for things like EBNF in a specification.
That said, I think you are correct that this will not take off as much as CommonMark has. I guess I was just sharing that I don't care and why.
[1]: https://gavinhoward.com/2020/04/i-am-moving-away-from-github...
On topic, though, preceding sublists with empty lines is a complete non-starter for me. However, since I don't hard-wrap lines (goal 7), but use soft-wrap only, I am not in the target audience anyways.
Both are clearly better than RST.
RE goal 7: here is my questioning of that goal and JGM's response, my counter, etc: https://talk.commonmark.org/t/djot-a-light-markup-language-b....
In other words, the plain text analog to WYSIWYG.
Professor MacFarlane must have really hated writing his papers in Latex.
> On topic, though, preceding sublists with empty lines is a complete non-starter for me.
I get it. This is my biggest complaint too. But I'm also a parser writer, and I see how this removes so much ambiguity.
This is the one thing that will stop djot adoption, and I'm okay with that. I don't think djot will be as used as Markdown, and I'm okay with that.
The odd newline requirements on lists and blocks, the special syntax for raw HTML and so on makes Djot feel more artificial to me.
My favorite number is probably the number
1. It's the smallest natural number that is
> 0. With pencils, though, I prefer a
# 2.
is a paragraph, a list item, a block quote, and a heading (4 things). In djot it's just a single paragraph.If you want it to be 4 things you have to add a newline between each one.
Personally I think djot gets this right.
- foo
- bar
- baz
- foobar
and demands this: - foo
- bar
- baz
- foobar
which is an unergonomic choice.This forcing new lines thing almost totally eliminates any desire I have to experiment with it, when I was very excited about it when I first saw it come up. Hopefully they'll change their mind. They want source code to be readable then make a decision that makes our eyes bleed.
So:
- list
- sublist
- sublist
is valid Markdown but invalid djot, while - list
- sublist
- sublist
is valid djot.Markdown is used enough that you are going to need to know the syntax. So a competitor doesn’t just have to better, it has to have enough additional merits to be worth learning in addition to Markdown.
It's easy, most projects can satisfice with it, and people on the projects that can't satisfice with it may not think about markup enough to realize they're painting themselves into a corner until they have a big ballast of existing documentation to cope with?
I've been fumbling around for how to convey signs that a project may need better tools (https://t-ravis.com/post/doc/what_color_is_your_markup/) but it's been slow-going and I'm bearish on how well ~better-practices will spread.
As for why it never gained the memetic popularity of Markdown that might have led to a different trajectory, that's harder to say. The One True Markdown is fundamentally much simpler than AsciiDoc, and consequently much easier to learn, easier implement in JavaScript for live rendering on the Web, and easier to extend with your own opinionated features. So I think it was easy and attractive for various platforms like Hacker News and Github to support it, and this I think had a snowballing network effect.
Personally I love AsciiDoc and I think it's the future of technical writing and publishing. It's everything I wanted out of reStructuredText but without its fussy, non-composable syntax. However I don't think that future will become reality until a spec is published that is friendly to implementers other than AsciiDoctor.
AsciiDoctor pretty much focused on a direct HTML translation and ignored the inconvenient parts. (Some of the inconvenient parts are deprecated syntax that while AsciiDoc's had a replacement for, I've written the old style for ~20 years and when GitHub tries to render a document with AsciiDoctor, oops; sometimes I'll change the document, sometimes I'll decide rendering on GitHub isn't important.)
I think the primary reason is because Markdown is prettier to read (and also secondarily because it's good enough).
There are lots of markup formats, but markdown is probably the easiest on the eyes.
The point on standardization is sort of moot until Djot becomes big enough to be more than a single implementation. I'd be happy to see it get to that level, but I won't be waiting for it either since Pandoc Markdown is perfectly servicable as it is.
I'm currently working on a second implementation, mostly because I want my docs to be dependency-free, but also for standardization of djot.
I mean you can do it - look at all the RFCs for example - but they must have been unpleasant to write and they're certainly unpleasant to read.
Sensible. Mostly since it makes other things easier (goal 5), second because one thing is only represented in one way, and thirdly (least important) since indented code blocks are kind of a pain to format compared to fenced code blocks.
On a practical level, nearly everything this does is better served by Pandoc Markdown today. But maybe in some not-so-distant future I'll be able to make use of this.
[1]: https://spec.commonmark.org/
[2]: https://talk.commonmark.org/t/issues-we-must-resolve-before-...
I wonder what someone like jgm has to say about the difference of writing Haskell and Lua.
I wonder if jgm’s djot parser being in Lua is also influenced by the fact that he is a Vim user:
(Assuming Neovim because this repo has some Lua files.)
So naturally, since the embedded lua is expanding its functionality over the years, including the ability to write custom writer and reader, this is the way to be able to tap into pandoc without committing. Tapping into pandoc is important here, as the community who likely will provide feedbacks and insights are there.
I'll bet eventually when it becomes mature, a native Haskell implementation will comes.
Shame that the creator of Markdown blocks any efforts to to fix or standardise the format.
https://blog.codinghorror.com/standard-markdown-is-now-commo...
Jeff Atwood goes out of his way to be courteous to Gruber in this post, but frankly, I think Gruber was being a jerk here, using his claim to the name to tyrannise an open source community that he has otherwise not been involved with in the slightest the last ~17 years.
If Github supported it alongside markdown, this would drive the markdown parsers to support this as a flavor, and then the community could decide; it will be fascinating to analyze the open-source population.
I'm finding dozens of applications for a lingua franca of mildly-structured, human-readable language that transforms as easily into data as documents and UI's. I've built out 10+ markdown applications using 3-4 parsers, and always had to reduce the effective semantics to a small subset of features, and then had to use metadata or convention to do what attributes will do much more cleanly.
As for Lua: yes it's the best for integrating into pandoc performance-wise, but ay! the table structures and metadata tables to get anything like typing are mind bending, and I am still miffed at being blocked at a critical time by a bug in pandoc's lua initializer for tables.
Aside language/stdlib developers, jgm is my greatest benefactor, and he deserves a cadre of loyal and competent implementors.
pandoc is interesting because with markdown it currently converts the "alt text" for figures  in markdown->pdf conversion into the figure captions, which is cool but makes me think: it is not alt anymore, anyone should have it
What we should have instead of markup, is a WYSIWYG with keyboard shortcuts. Confluence, for example, will convert Markdown into rich text in real time, and has keyboard shortcuts for its other layout/style options. But the point is to edit it in a GUI, see your changes live, and not need to learn a language in order to edit a document. There are so many problems you avoid by giving the user tools to make their life easier. Markup may be one tiny part of that, but it shouldn't be considered the complete solution.
Maybe I am misunderstanding what you are saying. Isn’t this the same as LibreOffice Writer and Microsoft Word? What is the existing problem that such an interface would solve?
https://prosemirror.net/examples/markdown/
From all web editors and their HTML transformations, ProseMirror has proven to be most sane for me.
I suspect I am not alone.
My biggest complaint is that asterisks map to <strong> and underscores map to <em> (in HTML terms). This is not backwards-compatible with Markdown where (asterisk)foo(asterisk) gets you <em>foo</em>, and it feels objectively backwards, if that makes sense. I wonder if there's any chance they could reverse that.
Interesting. I suspect something like this will always be subjective, but I find the opposite to be true. *bold* and _italics_ make the most sense to me and is always what I wished Markdown did.
Probably, this is because I was familiar with Textile[1] before I used Markdown, and this is what it does.
Today, Slack also uses this convention instead of the Markdown convention (though I believe it _used_ to use the latter).
Yeah… After posting that message, I remembered that _foo_ in Markdown also results in <em>foo</em> - so the underscores are backwards-compatible. But I've just always used asterisks so I completely forgot about it. So I guess they were bound to make some people upset no matter which one they made <em> and which one they made <strong>, and I'm on the losing side. :P
Whatever. If this ends up taking over the world as Markdown did (and I hope it does), I'll just get used to it, I suppose.
Incidentally, how did you "deactivate" HN's parsing of the asterisks in your reply?
Click the small “help” link on your own profile page: https://news.ycombinator.com/user?id=Cyberdog
The help link also shows if you re-edit a comment.
For example:
> Block-level elements can't interrupt paragraphs (or headings), because of goal 7
It then goes on to show they do interrupt paragraphs
- this then - this other thing
vs - this then
- this other thing
The 2nd is 2 list items but it's just the first with being interrupted by a block-level element.For me, I use Markdown because I can transform it with Pandoc and my own filters to any other kind of document:
I picked at what I mean a little in a post this summer: https://t-ravis.com/post/doc/the_gizmos_role_in_markup/
(I've also made the first post in an unfinished series that will continue to explore these ideas about markup, but a bit less directly: https://t-ravis.com/post/doc/what_color_is_your_markup/)
Similarly, for things like comments on message boards or blogs, a user can just dump a bunch of text into the text box without knowing the first thing about Markdown and expect it to look more or less like how it was entered, with paragraph breaks and such. If you force these people to use HTML instead, you're forcing them to at least learn and use <p></p> - which is probably simple for those of us reading HN, but I don't consider it a reasonable request for the normies on Reddit.
So, sure, HTML is quite good at what it was invented for, but not everything that involves text input on the internet or elsewhere should be HTML.
/dʒɑt/
/dʒ/ is like both the `j` and `dg` in english judge
/a/ is like the a in father in most American english dialects
/dʒɑt/ is also the complete IPA for the word "jot" (as in "to jot down") in American English