"Why not a parser generator"
Seemse like the way to go these days? Why write it yourselve if you can generate it from EBNF notation?
"Why not a parser generator"
Seemse like the way to go these days? Why write it yourselve if you can generate it from EBNF notation?
> [...] we are here to learn, we want to understand how parsers work. And it’s my opinion that the best way to do that is by getting our hands dirty and writing a parser ourselves. Also, I think it’s immense fun.
In other words: use a parser generator when you need a working parser, quick. But if you want to learn how parsers work, what ASTs are, what "top-down parsing" means, etc. then I recommend you write your own. It's not that hard once the concept "clicked" and, again, it's a ton of fun :)
I really like that philosophy - as Feynman said: "What I cannot build I do not understand"
This is true, but it provides no guidance on the question whether you need to understand it. If you don't want to build a web browser, you don't need to understand how web browsers are built. That doesn't mean that it is dishonorable for you to use a web browser.
So do you need to understand how generated parsers work? I think that's a matter of preference that Feynman doesn't help you decide.
Not necessarily. Use it also when you need a more correct one :-)
The way the book presents the parsing subject leaves the impression that a handcrafted parser is generally preferrable.
I have experience with a project that was based on that book, and there were (and probably, are) a lot of bugs in the parsing code, that wouldn't have been there if a parser generator was used.
I don't camp for using parser generator at all costs, however, presenting both sides of the coin would allow readers to make a (more) informed choice.
I think I need to a add a big "do. not. use. in. production!" disclaimer then :) The parser we build in the book is, of course, not a battle-tested, industrial-grade parser that can survive every fuzzer you throw at it. Far from it, but that's fine, since that was never the goal. The goal was always to learn.
So, yes, I agree. If you do need a production-ready, stable parser: use a mature parser generator or, maybe, invest more time in getting good at writing parsers than it takes to read and work through the ~100 pages I wrote on the subject.
> Why not a parser generator
IIRC it was about "minimising magic" - showing what's actually involved in native code. Bob Nystrom makes the same decision in "Crafting Interpreters [1]. Quoting [2]:
"Many other language books and language implementations use tools like Lex and Yacc, “compiler-compilers” to automatically generate some of the source files for an implementation from some higher level description. There are pros and cons to tools like those, and strong opinions—some might say religious convictions—on both sides.
We will abstain from using them here. I want to ensure there are no dark corners where magic and confusion can hide, so we’ll write everything by hand. As you’ll see, it’s not as bad as it sounds and it means you really will understand each line of code and how both interpreters work."
[0] https://corecursive.com/037-thorsten-ball-compilers/
[1] https://craftinginterpreters.com/
[2] https://craftinginterpreters.com/introduction.html#the-code
> There’s also some argument made by GCC and Clang teams that they can hand-tune a hand written- parser to get better error messages. I believe they can do such hand tuning and probably have. But that argument does not stop you from customizing the grammar for generated parser to do essentially the same thing, so I don’t buy it as a differentiator.
This is an excerpt from an interesting Quora reply on the subject: https://www.quora.com/Are-there-many-established-programming...
Ira Baxter makes a living out of selling proprietary technology based on GLR parsing. His opinion matters, but if there is no open source GLR parser generator available (is there? I don't know), the point he makes cannot apply to GCC or Clang.
https://www.gnu.org/software/bison/manual/html_node/GLR-Pars...
PROGRAM -> STATEMENT*
STATEMENT -> IF-STATEMENT | ASSIGN-STATEMENT | STATEMENT-ERROR
IF-STATEMENT -> ...
ASSIGN-STATEMENT -> ...
STATEMENT-ERROR -> TOKEN* SEMICOLON
The idea being that if the if-statement and assignment statement rules fail you consume tokens until the next statement separator (semicolon in our case), produce an Error node in the AST and resume parsing the next statement. This allows your parser to actually catch multiple errors in one pass since it won't die the moment it encounters something it can't parse. After parsing you scan the AST and if it has any Error nodes you spit out the relevant error messages and exit. You can get as granular with the error rules as you want, although the more you add the more unwieldy the grammar becomes.Naturally doing this in practice is harder than it sounds. Tweaking the Error rules is hard. If you consume too much or too little input you will resume in a spot that can have no chance of ever succeeding, giving way to a cascade of parse errors. I think we've probably all seen a compiler spit out a couple of dozen error messages all stemming from a single root error like a missing closing parentheses or brace.
Also it's questionable that they add very much. Parsing text isn't all that difficult. The hard part is to get the tree structure of the AST right.
All in all the parsing is a quite small part of a compiler. I've converted from a parser generated to manual code in two different projects and they actually ended up with less lines of code after the conversion, if just barely.
No I think people use parser generators less these days.
> Why write it yourselves if you can generate it from EBNF notation?
Can you write EBNF for it? Can you generate a parser from the EBNF? EBNF and most parser generators are designed for context-free languages. Most languages are context-sensitive. You can see where things start to go wrong...