Basically the key insight is that you can use the parsing automaton state to classify syntax errors, which makes it as easy as falling off a log to produce good messages. The paper originating this technique is Clinton Jeffery's TOPLAS 2003 paper “Generating LR Syntax Error Messages from Examples.” This is what go uses, and OCaml uses a more advanced version of this technique introduced by Francois Pottier in his CC 2016 paper "Reachability and Error Diagnosis in LR(1) Parsers".
https://github.com/golang/go/blob/master/src/go/parser/parse...
https://github.com/golang/go/blob/master/src/go/scanner/scan...
Why did you think it was generated?
Edit: I forked the compiler a while ago to add maybe types (a la Rust result). Looks like the compiler code has changed quite a bit from when I was playing w/ it.
Be sure that your language will parse. It seems stupid to sit down and start designing constructs and not worry how they will fit together. You can get a language that's difficult if not impossible to parse, not only for a computer, but for a person. I use YACC constantly as a check of all my language designs, but I very seldom use YACC in the implementation. I use it as a tester, to be sure that it's LR(1) ... because if a language is LR(1) it's more likely that a person can deal with it.
* https://github.com/antlr/antlr4/blob/master/doc/faq/general....
* See: "What do you think are the problems people will try to solve with ANTLR4?" question
I've used several different parser generators in the past. But I've also transitioned to hand-rolled recursive decent parsers, being Lazy I've created a library to assist in hand-rolling a recursive decent parser: https://github.com/SAP/chevrotain
[0] https://github.com/apache/spark/blob/master/sql/catalyst/src...
[1] https://github.com/prestodb/presto/blob/master/presto-parser...
[2] https://github.com/microsoft/TypeScript/blob/master/src/comp...
I wonder if this is relevant to how bad the error messages generally are in Typescript.
- Grammar documentation: https://golang.org/pkg/text/template/
- Code: https://golang.org/src/text/template/parse/lex.go
Edit: His name is Rob Pike
I think it does make sense to write manually parsers for performance and error messages, but it should be clear that this means raising cost by ~10 times. It is worth the effort if you are building a compiler for Java, for example, probably not if you want to process some DSL you developed.
Disclaimer: I am the brother of the author of this article
I have plenty. Publish contact info on your HN profile and I can send some.
https://github.com/cockroachdb/cockroach/tree/master/pkg/sql...
One of the more memorable parsers I’ve worked on was a parser for SPICE netlists. I started out believing that it wasn’t going to be too big of a deal, and ended up sinking a ton of time into it. Ultimately the company (as far as I know) ended up just buying a $40k license to some obscure company that had one, because there was a constant cat-and-mouse game of getting it working right and then discovering a customer who did even more strange stuff that was somehow accepted by the 3rd-party SPICE sim, but wouldn’t be accepted by ours.
I’m assuming that PHP has evolved a ton since I encountered this, but IIRC at one point you couldn’t do `$foo[$baz]($zap)` to call an anonymous function stored in an array, rather you had to `$t = $foo[$baz]; $t($zap)`. At the time (young and naive), I just couldn’t comprehend how a sane parser wouldn’t parse the first form, and then I started looking at how it was implemented...
I would argue that parser-generators are what made those projects to be started and prosper in the first place. Then once one is successful a custom solution could make sense but in my experience it is more expensive and potentially less maintanable, unless you know very welll your way around parsers and language tooling
To put it another way, I'm rarely parsing data for it to be directly optimized into a machine language. I'm parsing data to extract parts I care about and then work with those parts. The more consistent this process is the easier it is to debug and fix. If my whole parsing/using-the-parsed-data pipeline is in the same language, say Javascript, then I am still only ever debugging the same language and runtime environment (eg: node v12 on whatever \*nix).
In a practical example, YARA[0] is (confusingly) used as both a format[1] and specific implementation[2] for sharing malware detection rules. ClamAV[3] is a popular open source antivirus engine that added YARA-the-format support a few years ago[4]. If you look at their grammar[5] file as well, you can see that one uses GNU Bison 3.0.4 the other uses 3.0.5. One is 3754 lines long, the other is 1849. We can expect these to behave differently. At a certain point, say because of how regular expressions are handled[6], it becomes easier to maintain your own parser than to deal with quirks/whims of someone else's implementation (generated or not).
[0]: https://en.wikipedia.org/wiki/YARA
[1]: https://yara.readthedocs.io/en/latest/writingrules.html
[2]: https://github.com/VirusTotal/yara/blob/master/libyara/hex_g...
[4]: https://blog.clamav.net/2015/06/clamav-099b-meets-yara.html
[5]: https://github.com/Cisco-Talos/clamav-devel/blob/898c08f08b5...
[6]: "In previous versions of YARA, external libraries like PCRE and RE2 were used to perform regular expression matching, but starting with version 2.0 YARA uses its own regular expression engine. This new engine implements most features found in PCRE, except a few of them" from https://yara.readthedocs.io/en/latest/writingrules.html#regu... ; if the regular expression grammar and/or symbol set isn't consistent the things parsing the files won't necessarily be either