The base of the parser relies on a grammar to describe possible sentence structures, but actually most of the work in disambiguation is done using statistics, or with neural networks. The question what is and what is not grammar is rather arbitrary, there can be a continuum between simple rule-like and statistical regularities.
Ultimately the models in NLP are rarely making any kind of cognitive/neuroscientific claim about being plausible models, just effective ones. There's a specific field of computational psycholinguistics which does investigate those things.
> Also, it's really not the case that you get better performance out of neural network models of language, let alone "human like performance". We're very far from that still.
I wouldn't make such claims these days because deep learning methods are gaining ground very fast. There are many tasks at which the deep learning model is better than traditional models. Furthermore, there are already deep learning models which are better than humans at image labeling.