Similarly if you ask ChatGPT about the current president of Numbitistan it will tell you that it doesn't know about a county with that name, rather than just hallucinating an answer. So it can at least in this circumstance tell the difference between knowing something and not knowing something.
When robots are powered by transformers or the like, I expect we'll see some pretty impressive results.
But it's the internal consistency that really matters. LLMs have no internal consistency, because they have no way of 'adopting' a view, value, fact, or whatever else. They will randomly hallucinate things that directly contradict the overwhelming majority of their state, and then do so repeatedly in a single dialogue. If there were a human behaving in such a fashion, we would generally say they had schizophrenia or some other disorder that basically just ruins your ability to think like a human.
https://en.m.wikipedia.org/wiki/Compartmentalization_(psycho...
The biggest mistake I see people make when criticizing LLMs is that they take the best possible modes of human thought from our best thinkers, and compare that to LLM edge cases.
Accuracy vs consistency isn't really a delineator. There's so much low-hanging fruit atm, like world models for LLMs improving drastically if you just train them longer. I'll believe the naysayers if say in 5 years GPT-4 is still near state of the art. Until then, there doesn't seem to actually be any theoretical limitations.
Like, if you can't tell me "what day is it today?" (actual failed prompt I have seen) then there's no world where I'm going to have a more complicated follow-up conversation with you. It's just not worth my time or yours.
But has anything changed? Well no, because it's obviously trivially possible for them to play a decent game of chess (or correctly assess the date), but it's an example of a more general issue of LLMs being generally incapable of consistently engaging in simple tasks across arbitrary domains. So you have software that can score some high thing on the LSAT or whatever, but can't competently engage in a game children can play.
The over-specialization for the sake of generating headlines and over-fitting benchmarks is, IMO, not productive. At least not in terms of creating optimal systems. If the goal is to generate money, which I guess it is, then it must be considered productive.
But like, people are on here saying that this will make scientific improvements, and until it can get past the basic stuff, it's not in the ballpark of anything more complicated. Right now, we're basically at the stage of 10 million monkeys on 10 million typewriters for 10 million hours. Like, maybe we'll get Shakespeare out of it, but are we willing to sort through all of the crap it will generate along the way, when it can't actually create a useful answer to simple questions?
Take what I wrote above. If you were given the context of what I have already written, then you could probably fill in most of what I wrote, to a reasonable degree of accuracy, after "It is their normal..." Because the issue is obvious and so my argument largely writes itself. To some degree even this second paragraph does.