Chumsky: A Tutorial
github.com
github.com
... at tremendous cost to my sanity. I've later been told I may have inadvertently broken two of the seven seals.
Was considering what it would take to write a small ML that compiles to lua. "Almost certainly more than I'm willing to put in" was my tentative unresearched conclusion but this validates it a bit.
The lisp ones look solid and I love lisp, but for this specific thing a good type system would be most valuable.
1. Convert to continuation-passing style (or ANF)
2. Perform closure conversion
3. Hoisting nested functions/lambda lifting
4. Generate C! (and then write a runtime/GC).
Targeting lua should be much easier
Some resources I like:
- Andrew Kennedy's 2007 paper Compiling with Continuations, Continued [1]. This one is the most clear IMO
- Andrew Appel's Compiling with Continuations book (a bit outdated though... assembly code is for VAX)
- Matt Might's series [2]
- MLton's source and documentation [3]
[1] https://www.microsoft.com/en-us/research/wp-content/uploads/...
This is supposed to be a fun article for experimentation in programming languages, not a new advancement in a garbage collection algorithm for increasing the efficiency of a datacenter by 2%.
My favourite new programming language is Svelte for example (just a parser that emits javascript), it's just a little different from HTML/CSS/JS, but just enough to make web development fun again for me.
- Ease of use + ergonomics
- Backtracking 'by default'
- High quality errors and error recovery
Nom is very specifically designed for parsing input not intended to be read by humans, and that focus is evident in the things is prioritises.
But what this library and some others are aiming for is the NEXT level of parsing difficulty: Parsers that work AND also provide high-fidelity output for SEVERAL consumers (IDE, editors, linters, ...) AND provide high-quality errors (aka: diagnostics) AND provide suggestion to fix that errors.
That is another level!
MUCH HARDER!
> Whitespace is not automatically ignored. Chumsky is a general-purpose parsing library, and some languages care very much about the structure of whitespace, so Chumsky does too
TL;DR: The best approach, in my view, is to detect whitespace in the lexer and emit fake delimiter tokens for the parser so that the parser-level syntax is still context-free.
Maybe I’m being stubborn, but I prefer to write lexers and parsers by hand. Although, I would like to get around to writing a “common lexer and parser utility library” one day to abstract away some common patterns while simultaneously not forcing design decisions. Perhaps I’ll end up with a parser combinator library by the time I’m done and I’ll have to eat my words.
Most parser combinator libraries give you lots of help with the easy stuff (e.g. parse a list of integers separate by commas), but not so much help with the hard stuff like generating good error messages or avoiding unbounded backtracking.
Curious to find out more about critiques on his work though (in all spheres)