I'm not sure why you interpreted my comment as "antagonistic" or "unfriendly". Yes I do disagree with you (at least assuming I'm understanding you correctly), but I'm not attacking you personally, just your argument.
I guess it's fair to say my examples are cherry-picked too (though I didn't really give specific examples in my comment so much as entire general categories of problems ChatGPT is known to be proficient at solving). But they aren't so cherry-picked as for "random chance" or "that example was in the training data verbatim" to be possible explanations. It's not like I had ChatGPT answer 100 billion questions and am only showing you the top 0.1%. It's very common for ChatGPT to be perfectly correct even when answering logic puzzles or coding problems not in its training set. So if not those then what's your explanation other than "understanding"?
On the flip side, I don't think it's possible to prove ChatGPT does not posses understanding with individual examples. (And it seems you agree?) Those are easily dismissable as just "ChatGPT isn't great at that particular task". Particularly given how many other tasks ChatGPT is great at.
Scott Alexander figured this out back in 2019 before GPT-3 was even a thing (let alone ChatGPT), I think it's worth a read: https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
> The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer.
Not really though. The "pound of bricks or pound of feathers" question is specifically designed to trip up humans, and the specific formulation you used seems specifically designed to trip up LLMs (by playing off their tenancy to pattern match common sayings), yet despite those disadvantages GPT-4 succeeds.
Further, I don't think "I'd expect even a child to be able to answer this" is a good metric. ChatGPT isn't human, so we shouldn't expect it to be good at everything humans are good at just because it's good at some things humans are good at. Again, I'm not trying to argue ChatGPT is a human-like AGI.