This idea that GPT only works at the level of words and develops no deeper understanding of the concepts in language seems silly given its behaviour. And at the very least it's not what we observe from other NNs. As you point out a CNN will find deeper relationships and patterns between images, so it's only reasonable to assume a very large language model would find deeper relationships in text data.
The only difference here is that in comparison to other problems, text is how humans communicate and encode knowledge. The deeper relationships to be found in text is knowledge + reasoning.
I think we can say with some certainty that GPT models knowledge, the thing people are less sure about is if it learns to reason.
My take on this is that the fact you can ask it stuff that it couldn't know, but it can still "reason" to the correct answer suggests strong that it must have some ability to reason on the knowledge it's acquired.
Here's a really dumb example:
Me: Daisy likes to go swimming on the weekend, but last week she swore at her brother and has been grounded. How does Daisy feel?
GPT: It's possible that Daisy may be feeling disappointed or frustrated since she is unable to go swimming, which is an activity that she enjoys. She may also feel regretful or guilty for swearing at her brother and for the consequences that followed.
This isn't knowledge regurgitation. GPT doesn't know who is made up person is so it can't simply regurgitate something it was trained on. The only explanation for behaviour like this is that GPT has modelled human emotion and can reason about it.