Most of us when using LLMs feel it should only produce the one correct answer; sometimes that’s the right way to think about it, but other times we might want an opinionated answer, when there is no true answer.
So I think if we get better learning and self correcting machines, we also get more opinionated and sometimes incorrect machines. Kind of like people. And sometimes they’re stubborn or confidently incorrect, but that I think actually points to more intelligence (in people at least), since it points to people deciding based on their experiences.