I've discussed related topics but haven't had the chance to address it directly. I'm glad someone is paying attention :)
Concretely, the first two links in this post show (old versions of) frontend/lex.py and osh-lex.re2c.h. TODO for me: put up the latest versions, as well as the huge C file with state machines that re2c eventually generates.
When Are Lexer Modes Useful? http://www.oilshell.org/blog/2017/12/17.html
It works like this:
( ) frontend/lex.py (a bunch of Python regexes, has some "metaprogramming") ->
(+) frontend/lex_gen.py ->
( ) osh-lex.re2c.h (re2c input file) ->
(+) re2c ->
( ) osh-lex.h (state machines in C, i.e. DFA as a big switch/goto)
where the (+) nodes are compilers, and the ( ) are source code files.
The reason I call this "a mathematical dialect" is because the same regular expressions run under Python's re engine (a backtracking engine) and as native code via re2c, an automata-based compiler.
If you scroll toward the bottom of this doc there's a useful table:
Regex Theory and Practice http://www.oilshell.org/share/05-31-pres.html
One side is "Perl-style regexes", which Python's engine is based on. The other side is "regular languages". Regular expressions were always mathematical, but the name got taken over by programmers to mean something different, so I call them "regular languages" or this "mathematical dialect".
These articles explain the difference,
https://swtch.com/~rsc/regexp/
but they're very long and a lot of people still don't understand the difference.
https://news.ycombinator.com/item?id=20311630
(which is understandable since it mainly comes up in performance corner cases, and when you want to compile regexes, which most people don't do)
I should write about this, but the lexer is one of the more solid pieces of the project. That is, it's "done" for now, and I need help with all the other parts, so I prioritize writing about those parts!!!
More lexing posts: http://www.oilshell.org/blog/tags.html?tag=lexing#lexing