The problem is a bit deeper than that, because what we perceive as "confidence" is itself also an illusion.
The (real) algorithm takes documents and makes them longer, and some humans configured a document that looks like a conversation between "User" and "AssistantBot", and they also wrote some code to act-out things that look like dialogue for one of the characters. The (real) trait of confidence involves next-token statistics.
In contrast, the character named AssistantBot is "overconfident" in exactly the same sense that a character named Count Dracula is "immortal", "brooding", or "fearful" of garlic, crucifixes, and sunlight. Fictional traits we perceive on fictional characters from reading text.
Yes, we can set up a script where the narrator periodically re-describes AssistantBot as careful and cautious, and that might help a bit with stopping humans from over-trusting the story they are being read. But trying to ensure logical conclusions arise from cautious reasoning is... well, indirect at best, much like trying to make it better at math by narrating "AssistantBot was good at math and diligent at checking the numbers."
> Hallucinating
P.S.: "Hallucinations" and prompt-injection are non-ironic examples of "it's not a bug, it's a feature". There's no minor magic incantation that'll permanently banish them without damaging how it all works.
Say, they should be 100% confident that "0.3" follows "0.2 + 0.1 =", but a lot of floating point examples on the internet make them less confident.
On a much more nuanced problem, "0.30000000000000004" may get more and more confidence.
This is what makes them "hallucinate", did I get it wrong? (in other words, am I hallucinating myself? :) )
Overconfident people ofc do not contribute positively to the system, but they skew the system reward's calculation towards them: I swear I've done that work in that direction, where's my reward ?
In a sense, they are extremely successful: they managed to do very low effort, get very high reward, help themselves like all of us but at a much better profit margin, by sacrificing a system that, let's be honest, none of us care about really.
Your problem maybe, is that you swallowed the little BS the system fed you while incentivizing you: that the system matters more than yourself, at least at a greater extent than healthy ?
And you see the same thing with AI: these things convince people so deeply of their intelligence that it blew to such proportion that NVidia is now worth trillions. I had a colleague mumbling yesterday that his wife now speaks more with ChatGPT than him. Overconfidence is a positive attribute... for oneself.
If one contributes "positively" to the system, everyone's value increases and the solution becomes more homogenized. Once the system is homogenized enough, it becomes vulnerable to adversity from an outside force.
If the system is not harmonious/non-homoginized, the attacker would be drawn to the most powerful point in the system.
Overconfident people aren't evil, they're simply stressing the system to make sure it can handle adversity from an outside force. They're saying: "listen, I'm going to take what you have, and you should be so happy that's all I'm taking."
So I think overconfidence is a positive attribute for the system as well as for the overconfident individual. It's not a positive attribute for the local parties getting run over by the overconfident individual.
Of course, the result is that people get fed up and decide that the problem has been not that democratic societies are hard to govern by design (they have to reflect the disparate desires of countless people) but that the executive was too weak. They get behind whatever candidate is charismatic enough to convince them that they will govern the way the people already thought the previous executives were governing, just badly. The result is an incompetent tyrant.