Ohm: A library and language for building parsers, interpreters, compilers, etc
github.com
github.com
We've been on here a few times before:
• Ohm – A library and language for building parsers, interpreters, compilers, etc.: https://news.ycombinator.com/item?id=26603393 (March 2021)
• Ohm - Parsing Made Easy: https://news.ycombinator.com/item?id=15491336 (Oct 2017)
You might also want to check out WebAssembly from the Ground Up, an online book we're writing that uses Ohm to teach you WebAssembly: https://wasmgroundup.com
More specifically, can I build lexers with Ohm? Can it generate a syntax diagram from a grammar?
Ohm is focused on being easy to use. Grammars are generally very readable, and people love our online editor (https://ohmjs.org/editor/). Performance is fine for many use cases (hobby programming languages, query and schema parsers, etc.) but not good enough for a production programming language.
Chevrotain is much more focused on performance — honestly, it blows Ohm (and some other tools like ANTLR and PEG.js) out of the water. The tradeoff is that writing grammars is a bit more complex than with Ohm, and the result is not as readable.
I wonder if it can also deal with ambigious grammars and/or if it is a back-tracking parser.
It also has a single grammar for lexical and syntax description. (My interpretting parser is based on a hard-coded lexer.)
[1] https://fransfaase.github.io/ParserWorkshop/Online_inter_par...
The current implementation does use a tree-walk interpreter, but I'm considering creating a version that compiles to WebAssembly.
Truth be told, I'm glad I didn't know about it when I wrote a much more simplified project (shameless plug: https://github.com/catapart/Magnit.Tokenization), because I DEFINITELY would have just used your solution, even though its a bit overkill for those needs.
That said, after having finished what I needed, of course I started to wonder about what else I could add to it, with the main stopping force being the need to rewrite the parsing engine (regex ain't going to cut it for more complicated syntaxes). Which is one of those dev projects that linger in the back of your mind until you either see it through, or see that someone else has done it.
And, on that record, I think you've done a better job than I could ever attempt, so I'm very glad to know about this library, now! I don't have anything specifically in mind for it, but having the doors it opens available is quite nice!
The examples have a lisp-like interpreter at https://github.com/ohmjs/ohm/blob/main/examples/simple-lisp/... which definitely uses a grammar for parsing and might use a generic AST representation.
Will have to think more - a grammar might be a worthwhile way to specify a nanopass style compiler pipeline.
Usually they never end up in production, as while quite productive to prototype, they aren't that good in error recovery and meaningfull error messages.
I once used Codemod [0] to migrate an old JS codebase. Would this be a use case for Ohm as well?
I'd also love to hear of any existing tooling that does this
I have used both Ohm and Lezer - for different use cases - and have been happy with both.
If you want a parser that makes it possible for code editors to provide syntax highlighting, code folding, and other such features, Tree-sitter and Lezer work great for that use case. They are incremental, so it's possible to parse the file every time a new character is added. Also, for the editor use case it is essential that they can produce some kind of parse tree even if there are syntax errors.
I wouldn't try to build a syntax highlighter on top of Ohm. Ohm is, as the title says, meant for building parsers, interpreters, compilers, etc. And for those use cases, I think Ohm is easier to build upon than Lezer is.
For practical purposes, PEGs are always linear time and can easily handle prioritized choice, with the downside that declared priority is the only way to resolve ambiguity. Backtracking is cheap but they are stateless in a very specific technical sense, so can't handle some semantic constructs. The no ambiguity thing means they only succeed or fail, no incremental parsing of incomplete grammar for example like you would want in a syntax highlighter. Some PEG libraries aren't "pure" PEGs and can get around these limitations though.
In real use they tend to have a very straightforward mapping between declaration and grammar. I got used to them in Janet, which has them in the standard lib instead of regex and for that purpose they are incredible. Much easier to write and especially edit. They also work well for some languages but there are some real world language features that they simply can't parse.