I assume the lispy type languages fall into the easily parsed bucket, perhaps tcl as well? And Perl may be a good example of a language that has notable ambiguity?
I assume the lispy type languages fall into the easily parsed bucket, perhaps tcl as well? And Perl may be a good example of a language that has notable ambiguity?
Python, Java. Python generates a top-down parser from Grammar/Grammar using its own pgen.c system, and the Java spec has a grammr. This doesn't mean there's no ambiguity, but they are easier.
* Pretty easy:
JavaScript, Go. I think the semi-colon insertion rule kind of breaks the lexer/grammar model, but they are still easy to parse.
* Hard:
- C (http://eli.thegreenplace.net/2007/11/24/the-context-sensitivity-of-cs-grammar)
- Ruby (related thread: https://news.ycombinator.com/item?id=13825975)
- OCaml -- AFAIK, a naive yacc-style grammar will generate tons of shift-reduce conflicts
- Haskell -- based on hearsay, I think this is hard
* Insane:C++, Perl, to some extent bash. See the end of my post "Parsing Bash is Undecidable" for links: http://www.oilshell.org/blog/2016/10/20.html
I'm not as familiar with languages like PHP and C#, but they're somewhere between the easy and hard categories.
Pascal is also pretty clean and regular. Unlike Smalltalk, it was designed more to make the compiler's job easier, so not only is the grammar pretty regular, but it's organized such that you can parse and compile to native code in a single pass.
C is a pain to parse, it's context-sensitive in some areas. C++ inherits all of that and adds a lot more on top.
My understanding is that Perl can't be parsed without being able to execute Perl code since there are some features that let you extend the grammar in Perl itself?
Lisp seems like it should be nearly as simple: recursively read in a balanced s-expression and attempt to intern the atoms in that s-expression; error out if any of those steps fails. However, Common Lisp's scanner/parser (called the "reader") allows the user to define arbitrary macros; basically the scanner has access to the full compiler if you want it to, so it can get really complicated really quickly.
C is not LR(1) for annoying reasons (typedef). Same for Go, IIRC. Otherwise, it's okayish. Javascript has this inane "optional semicolon" thing, which probably makes it annoying.
C++ is probably the most insane thing to parse. And then, there is bash[1].
r"…", r#"…"#, r##"…"##
br"…", br#"…"#, br##"…"##
you need to count the #s.