I have this "meta-language" rant brewing for my blog[1], and I feel like you took the words out of my mouth!
Are you saying that ANTLR makes some unwarranted assumptions about Unicode? Does that just depend on the generated Java code, and would it be different for generated parsers in C++? I don't know the JVM very well, but my understanding was that the JVM makes some Unicode assumptions that aren't always appropriate.
The last "rant" was here, about meta-languages only being suitable for toys: https://news.ycombinator.com/item?id=13040682
vidarh also responds with his problems parsing Ruby (a "real" language): http://hokstad.com/compiler
The blog is based on my experience in writing a parser for bash. I ported the POSIX shell grammar to ANTLR, but it fell ridiculously short of being a production quality parser.
I believe it's essentially impossible to write a bash parser with ANTLR. But I don't hear about any alternatives. All I heard is yacc/bison for bottom up parsers, and ANTLR for top-down. Shell needs a top-down parser because it's an interactive language (the PS2 prompt) and for completion. But it doesn't appear there is any other "serious" meta-language for top-down parsers other than ANTLR? It certainly is the best documented and longest-lived, but I am surprised how far short it falls for many tasks.
Part of the rant will be a survey of "real" language parsers (i.e. take the top 20 TIOBE language implementations) and see that they almost all use hand-written parsers, or bespoke parser generators. For example, Python has its own top-down parser generator in the tree, pgen.c.
[1] http://www.oilshell.org/blog/2016/11/20.html