Let us construct IceCreamGPT. We take a corpus of text written by people who like ice cream and have provably demonstrated their joy while eating it. We then fine tune GPT 3.5 and the resulting model is called IceCreamGPT. Does IceCreamGPT like ice cream or is it only seemingly liking ice cream? It obviously likes ice cream, since it shares the same intentionality as humans responsible for the training data.
Now do the same with people who don't like ice cream but lie and write that they like ice cream. The performance of the second model is identical to the first model. Does this mean IceCreamGPT2 likes ice cream? Of course not, IceCreamGPT2 doesn't like ice cream despite it saying it likes ice cream! We know it doesn't like ice cream because it has the same intentionality as the humans responsible for its training data.
Now we have entered a magic world in which anything can mean anything.