Re2c: A free and open-source lexer generator for C and C++
re2c.org
re2c.org
https://sourceforge.net/p/joe-editor/mercurial/ci/default/tr...
It uses a variant a Thompson's matcher, see:
https://swtch.com/~rsc/regexp/regexp1.html
"However, Thompson-style algorithms can be adapted to track submatch boundaries without giving up efficient performance. The Eighth Edition Unix regexp(3) library implemented such an algorithm as early as 1985, though as explained below, it was not very widely used or even noticed. "
So in JOE's matcher, when simulating the NFA with the multithreaded machine, each succeeding thread will correspond to one of the sub-match possibilities. Augment this machine to keep the ones you want (the ones with the longest matching strings starting from the left usually).
Yeah, this is the standard way to do it. It's what I meant by "fall back to a slower engine," since simulating the NFA is much slower than a DFA.
https://github.com/ninja-build/ninja/blob/master/src/lexer.i...
and then used it for my own huge shell lexer:
The Oil Lexer: Introduction and Recap http://www.oilshell.org/blog/2017/12/15.html
The nice thing about it is that it doesn't impose any constraints on the structure of your code.
lex/flex probably have options not generate piles of messy C code with global variables at this point, but IMO there's no reason not to use re2c instead, so it's moot.
https://github.com/cjohansson/emacs-phps-mode/blob/master/ph...
For example, my JSON parser that doesn't use intermediate object tree and works directly on strings: https://megous.com/git/sjson/tree/sjson.c
re = regex.compilemulti([
(".*@(.*)\\.(com|org|net)", TagEmail),
("http://.*, TagURL),
("[a-z]*", TagWord)
])
match regex.tagmatch(re, pat)
| (TagEmail, submatches): domain=submatches[1]
| (TagURL, submatches): url=submatches[0]
| (TagWord, _): /* nothing */
;;
I haven't had time to do it in Myrddin's libregex (https://myrlang.org/doc/libregex/) yet, but I've got plans to do something like this. regex1 { code that executes if regex1 matches }
regex2 { code that executes if regex2 matches }
regex3 { code that executes if regex3 matches }Like... re2c?
(nice with https://github.com/fbb-git/bisoncpp as parser generator)
Very well written with good documentation.