We aren't testing whether the model's results are stable or correct for a given class of problem. The goal is to establish whether the model can reason.
Nothing capable of reasoning would contradict itself so blatantly and in such a short span while failing to indicate any kind of uncertainty.