I think what it means to understand language is to be able to generate and react to language to accomplish a wide range of goals.
Neural nets are clearly not capable of understanding language as well as humans by this definition, but they're somewhere on the spectrum between rocks and humans, and getting closer to the human side every day.
I can't help but think that arguments that algorithms simply don't or can't understand language at all are appealing to some kind of Cartesian dualism that separates entities that have mind from those are merely mechanical. If that's your metaphysics, then you're going to continue to find that no particular mechanical system really understands language, all the way up to (and maybe beyond) the point where mechanical systems can use and react to language in fully all situations that humans can.
What is understanding? Are we sure human understanding is not just patterns and combinations of patterns in our huge volumes of experience?
In this case, these models have vastly different architectures, learn language is vastly different format, and are given massive corpora of words written from millions of random perspectives instead of receiving words from one perspective via targeted communications.
On what basis is there to expect such a system to magically develop the expressive intentional characteristics of human language when we can't even find another animal who seems to have this ability?
There are a number of facts suggestive of this viewpoint:
For one, the apparent generality of the algorithms, resulting in a great deal of neuroplasticity and data-agnosticism. For example, you can wire visual input into auditory cortex and that auditory cortex will learn to see. [1]
For two, the similarity of Gabor filters learned automatically by statistical inference algorithms and the response properties of visual neurons. [2]
For three, the fantastic success, relative to other techniques at least, of such algorithms on natural language tasks such as machine translation.
Even its deficits are suggestive, such as the fact that the neocortex, like most statistical inference algorithms, can't easily do one-shot learning — for that you need a hippocampus.
Present-day machines obviously lack many abilities which humans have. Until we have a fully-functional AGI, it's hard to know for sure that there isn't a major missing piece which is nothing like the algorithms now under development. However, I'm not aware of any other algorithmic hypotheses of how the mind works which seem to so naturally fit the neuroscientific facts.
[1] Sur et al (1988). Experimentally Induced Visual Projections into Auditory Thalamus and Cortex. https://doi.org/10.1126/science.2462279
[2] Olshausen and Field (1996). Emergence of simple-cell receptive field properties by learning a sparse code for natural images. https://doi.org/10.1038/381607a0
The meaning of human language is to do with having an effect on the world, having a perception of the world, and having motivations related to the world - such as conquest, subjugation, hearing lamentations etc.
It seems possible to have the information content of language without the above extension in the world. Information is "just" combinations of patterns.
Humans need surprisingly little data to learn language. If we can discover our internal presumptions (a priori probabilities), AI mightn't need so much data either.
To use a crude metaphor, language could be just an API schema. Is there content accessible via the schema that is more meaningful than the schema itself? For APIs, yes. I think so for language too.
Not to mention the possibility of side channels. That, at the same moment any human reads any text, they receive extra information to help them understand that information.
And yet, if you ask them what the meaning of a given text is, they’re getting disturbingly close.
What they don’t have yet is self-awareness, a model of their own repeated interfacing with the world. And an understanding of language in that context.
But it’s coming.
Sorry but can you be a bit more specific here? How exactly has NLU split from NLP?