Do you know that the answer key wasn't part of the train set of ChatGPT?
Why should we assume that the training set includes the answer key in the absence of evidence?
It's significantly more likely that a model outputting the right answer was trained in that answer than not