Thanks for the response. Yes I think we are agreeing -- I have 2 of your books, and have read many of your ANTLR-related papers, and they were definitely the thing that taught me the most about top-down parsing. (I also used to share an office with Guido van Rossum and I remember he had your books too.)
I ported the POSIX shell grammar to both ANTLR v3 and v4 (which was basically changing yacc-style BNF to EBNF). But as mentioned, I discovered that the grammar only covers about 1/4 of the language. bash generates code with the same grammar using yacc, but fills in the rest with hand-written code. Every other shell I've encountered uses a hand-written parser. bash says they regret using yacc here:
http://www.aosabook.org/en/bash.html
I agree with you that there is a Pareto or long tail distribution in parser use cases. Most languages CAN use something like ANTLR or bison. But the parent was making a different claim:
What an odd statement, in light of the innumerable deployments of Bison / ANTLR parsers you certainly use at least once a day (if you spend any time at all in a terminal).
I would say that is FALSE, because most parsers that your fingers pass through are HAND-WRITTEN, because of the Pareto distribution. 99% of anyone's usage is of probably a dozen or so parsers, and they are either hand-written or generated by custom code generators, not general-purpose code generators like ANTLR or yacc.
-----
As feedback from a user of parsing tools, you might also be interested in my article here:
https://news.ycombinator.com/item?id=13628412
Someone is asking if there are any parsing tools that generate a "lossless syntax tree".
Also, based on my experience with ANTLR v3 vs. v4, I ask the question why use a concrete syntax tree at all? Nobody answered that question in the comments. I don't understand why that is a good representation, other than the fact that you might not want to clutter your grammar with semantic actions ("pure declarative syntax").
To me the parse tree / CST seems to be resource-heavy while containing unnecessary information, and also lacking some crucial information like where there's whitespace and comments.
To summarize my article, I'm researching code representations in the wild for both style-preserving source translation (like go fix, lib2to3 in Python) and auto-formatting (like go fmt).
It's definitely possible I misunderstood something since my experience was relatively limited, but I have read a lot of the docs and bought the books.