TatSu takes grammars in variation of EBNF, outputs memoizing Python PEG parsers
github.com
github.com
But now, I discovered LALRPOP[0] and Logos[1] in Rust which is just so much more powerful, I'm using PyO3 to bring it to Python because I find it easier to walk the AST and do code generation from here.
In Java (and maybe C++) ANTLR is a go-to option. Some languages have their own goto options, and others are lacking. When I did a quick research, Python seemed to be on the "lacking" side (multiple not-very-well-maintained-options but no clear winner). Is that changing (or was I missing something)?
https://github.com/lark-parser/lark
https://github.com/neogeny/TatSu
Lark seems to be the one with traction.
ANTLR does have python as one of the output languages IIRC.
There’s also a flex/bison clone written in python whose name I also can’t remember.
My current favorite parser generator library, Coco/R, has a python 2 implementation which I’ll maybe get around to porting to py3k one of these days, who knows?
But to answer your question I don’t think there’s a go to library on the python side, I’ve played around with a bunch of them and usually just go for re2c/lemon or Coco/R from C(++) since I mainly just poke at these things probably more than I should.
The next time I get a few days with nothing to do I plan to see if I can’t whip up a working lpeg (lua’s peg parser VM library thing) in python. Someone did a port a while ago but it is abandoned and, honestly, not all that good. No offense to the person who did it but you can do some pretty nifty things with the python C-api with some practice.
The flex/bison clone you're perhaps thinking of is PLY from David Beazley.
> The original version of this was written by John Aycock for his Ph.d thesis and was described in his 1998 paper: “Compiling Little Languages in Python” at the 7th International Python
http://pages.cpsc.ucalgary.ca/~aycock/spark/ (Older)
https://pypi.org/project/spark-parser/ (More recent)
It pertains to counting duplicate nested blocks as well as enforcing 1+, 1:M and N:M combinatorial of syntax block.
I once did the entire ISC Bind9 named.conf parser into PyParsing (named is a handrolled parser, not a Bison/Flex).
But it cannot do N:M nor enforcing 1:M.
So, can TatSu help with that AST-output part?
Tree Sitter (I think) does that but (I believe) its output is a parse tree.
For the N:M stuff, if I understand you correctly, sounds like global value numbering could do that as, if I understand it correctly, it gives you subexpression deduplication for free. I think this is a tree transformation step though — haven’t really thought about it before but it could probably be implemented as the parser output but would be destructive as you’d lose line number information and whatnot.
Personally, for my playing around, I use asdl to generate the AST nodes because its a super-simple ‘language’ so I can easily test out new things. I’ve reimplemented it probably five or six times now but that definitely qualifies as yak shaving.
Yak shaving. Sigh. It is still a thing at the frontier of leading-edge technology.
Will look at TreeSitter (again) and adsl.
That reminds me, I’ve been relooking at things from time to time and need to do a summary recap on paper. Blog it, perhaps.
I would love to hear more about how it affects cache hits.
Please submit an issue demonstrating it if you can! I'll do my best to fix it.