It's possible that the internal version of protoc is very different from the open-source version. (I know there are numerous differences, but not sure how pervasive they are in the parser.)
The open-source version has a hand-written tokenizer and recursive descent parser that is not too difficult to translate to EBNF. You'll notice that the section on numeric literals is a little wonky, because the tokenizer does a check that is hard to describe in EBNF. But it isn't too bad.
Also, some of the constraints of the language are in prose in this spec because they are easier to enforce using a semantic validation pass, instead of trying to model purely with a CFG. (Optionality of the colon in the text format, used in message literals, comes to mind.)
There are some things that technically _could_ be handled in the grammar, but they would make the grammar much more cumbersome to read and understand. So those things are also extracted into prose.
> Definitely interesting for this company to create an EBNF definition for protobuf.
For what it's worth, Google has also published an EBNF definition (the subject blog post contains links to those specs). But they are incomplete and not entirely accurate, which is a non-trivial part of what led us to writing and publishing this spec.