But I don’t think this is a nitpick at all. GPT models hallucinate information. They are right surprisingly often, but they’re also wrong quite often too. And the problem is they are just as confident in either case.
This is a fundamental, irreconcilable issue with statistical language models. They have no grounding in auditable facts. They can memorize and generate in very plausible ways but they don’t seem to have a concrete model of the world.
Ask ChatGPT to play chess. It can generate a text based board and prompt you for moves, but it can’t reliably update its board correctly or even find legal moves. Note that I don’t expect it to play good moves, but the fact that it can’t even play legal moves should tell us something about its internal state.
Now that GPT3 has trained on the whole internet, we may have reached a practical limit to how far you can get by simply training on more data with 1 or 2 orders of magnitude more parameters. There’s only so far you can get by memorizing the textbook.
At a more practical level, for most professions “pretty good” isn’t good enough. It’s not good enough to have code that’s right 90% of the time but broken (or worse, has subtle bugs) the rest of the time.