(1)
You realise that we don't investigate human understanding by looking at input-ouptut text correspondances? (This is quite a funny claim!)
The reason we think people have the capacity to interpret sentences is their context-appropriate sensory-motor actions, reasoning, emotion, etc.
When I say, "quick! a fire!" people: panic, run, get organised, etc. There are a very very large number of highly complex environmental, interpersonal, sensory-motor, emotional, etc. etc. interactions which follow.
ChatGPT not only does not display any, but cannot. This isn't up for debate, the system has no capacity to understand text. If I say, "quick! a fire!" it will not do anything which accords with showing an understanding of that.
Worse, as a matter of fact, we know its possible to generate output as-if it did understand. So we (1) have known a mechanism whereby it can fake it; and (2) know that this mechanism is the one its using.
(2) All thinking possess intentionality which simply means its about something, and the reason we ask, say, "should i call the police?!" is because we are situated in an environment, which causes us to form mental states about that environment.
Any system which emits, "should I call the police?" without having a backing representational state of, eg., the police, care, concern, priority, the environment which requires police, etc. *Does Not Mean* this question. It cannot.
(3) The capacity to speak truthfully does not mean you speak truthfully. That capacity is simply that you can form representational states which you can introspect as being accurate or inaccurate. ChatGPT lacks that capacity.
It does not form representational (intentional) states when it generates text; each sentence is not caused by such a state.
(4) The capacity i named was "caring", pretty much all people care, it's necessary for goal-directed action.