You wouldn't catch an early syntax error and would go on tokenizing till the end for nothing.
You wouldn't catch an early syntax error and would go on tokenizing till the end for nothing.
Most compilers try not to fail when they encounter a syntax error, they try to "recover" and parse the remaining document, usually starting from the next valid statement or declaration. This lets them report more syntax errors and moreover, is very important for good IDE support (if you've ever used an IDE/language combo where you only get completion after you've written the code, you know why). Although if the syntax error messes up tokenization (e.g. missing quote or end comment) it usually screws up the rest of the document anyways.
And some don't even generate an AST. :) Just read in and emit or interpret.
> Some compilers tokenize while parsing, but for a different reason: it's faster and uses less memory
Rather legit reasons.. The one I'm writing does this. It seems akin to natural language processing. You'd interrupt a speaker early if you can't make sense of his uttering. > Most compilers try not to fail when they encounter a syntax error, they try to "recover" and parse the remaining document
This seems ardous for you'd maybe have to keep parallel explanations of what you read ?But if you're reading a book and can't understand a sentence, you'll probably glance at the rest of the page to see if there are clues in the context.
Maybe it's a printing error and a paragraph has been repeated. A human who sees the entire page will detect that problem immediately and simply skip over the repeated paragraph. A computer that gives up at the first sight of incongruency has no idea that recovery was so easy.
The analogy here might be relevant to error messages. A compiler that "sees the whole page" can potentially offer more useful suggestions to the programmer about how to fix a problem.
In general it's useful for a compiler to not just hit the first error it can find in the source code and immediately error out. Instead, if you keep parsing after encountering an error, you can often encounter more errors, so that you can give the user a list of errors they need to fix, not just the first one. In that sense, a compiler's job isn't over the instant it encounters a syntax error, so the extra tokenization would not be useless.
String s = hello world";
[-> typename:'String' id's' op'=' id'hello' id'world' ...]
And it would cascade if then foo("blah"); (...)
[-> strlit';\nfoo(' id'blah' ...]
A parser-first approach will stop with unknown id 'hello'
and avoid the cascade.Users love consolidated, efficient feedback without having to peel back errors one at a time. It’s a distinguishing feature when you can find a way to do it that suits your input, and specifically so because it can be a hard problem to solve well!
1. Hand-written, recursive descent
2. Error matching expanded grammar beyond minimal gramar
That's interesting! I always assumed it was always more complex than that, I love it. How reliable is this method btw? I'm specifically curious about the risk of catching false positives due to failing to parse some sections, such as missing goto labels, or unused variables.
is rather a very good reason to interleave tokenizing and parsing. Otherwise your tokenizer will just get stuck in iowait for no good reason.
The other comments here about catching all errors is valid, but you don't need to store any of the previous output to do so. The tokeniser can provide a token at a time as necessary, and thus also provide errors in the same way.
Thinking otherwise just leads to suffering. The problem with tutorials like Tiny Compiler or Crafting Interpreters is that the authors do not run into problems with the code shown in their teaching materials, but as soon as a student wants to apply it to modestly complex grammars, it stops working. The authors traded conciseness for correctness which IMO is a bad trade-off, especially since a reasonably complete implementation of a parsing algorithm that has no shortcomings is perhaps only four to six times longer than a short and flawed one.
The point of this critique is to raise awareness; each student should not have to figure out that they got the bad end of said trade-off by trial and error, instead the teaching material should make this clear initially.
>> never
> nope you are wrong. I never tokenise everything up front
Explain how you could possibly arrive at the complete opposite meaning of what I wrote when you even use the same word "never" as I did?
Quick, get another downvote in, that'll teach me! Peak HN moment here.
I upvoted your correction.
Better off pulling tokens from semantic analysis through parsing top down, LALR(1), SLR(1), or alternative context-sensitive parsing. Derivation parsing is scannerless.
Fast lexer + cheap validator is a winning combo.