>> Whatever is behind my surprise, I am confident that the scientific method will lead to explanations.
I agree! Wholeheartedly so. But for the time being, the scientific method is not being applied. All that's been done is willy-nilly poking of different models and ooh'ing and aaah'ing at what falls off.
Of course, scientific explanations must take into account existing knowledge, the knowledge encoded in accepted scientific theories. We don't have any scientific theory of "understanding", in humans or machines. What we do have is a very clear theoretical and practical knowledge of how language models work. They are machines (in the abstract sense) that estimate the probabilities of sequences of tokens. Any explanation that fails to take this knowledge about what a language model is, and lack of knowledge about what "understanding" is, into account, will have to do a great, big deal of work to present a new theory.
And I would like to see such a new theory. In particular, perhaps we could have a theory of "understanding" in machines, based on current observations of the behaviour of large language models, and the known principles of their design.
But, so far, we have nothing like that! We have hand waving, wild proclamations based on faith and nothing else. It's impossible to reason for or against matters of faith.
>> Sure, it is unsurprising that a language model can calculate the probability of a string in a natural language without understanding language, but what I find surprising about it is that this alone quite often results in responses that could pass as human-generated.
I don't find that surprising. There are plenty of examples of systems capable of interacting with humans by generating natural language responses that "could pass as human-generated". For a couple of famous examples, SHRDLU, ELIZA and Eugene Goostman; they should be easy to search for online, otherwise please ask me for links. We know very well by now that this is no way to figure out the capabilities of a system, comparing it to human behaviour. That is particularly so for systems that are specifically created to mimic human behaviour.
You see, that's the big problem we're neck-deep into. Language models are machines that mimic human language production. By observing how good such a system is at producing human-like language, all we can say is how good the machine is at what it's designed to do. We can't draw any other conclusions. Not safely, because there is a great, major, risk of confirmation bias, and of circular reasoning, waiting in the wings. That would be so for any system designed to mimic human behaviour, but for a system that mimics language that's even more so, because it is extremely difficult to disentangle grammatical text from the expectation that it was written using human faculties.
tl;dr: we 're in a bias pit and we'll keep falling down it until someone figures out how to measure the abilities of LLMs somehow else than just poking them.