As a Human, I do not need to know it is a Language Model.
As a Human, I do not need to know it is a Language Model.
This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).
I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.
Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.
If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.
We don’t really need to speculate, this is already a problem.
I was unintentionally being pedantic, because this isn't really done with RL anymore - it doesn't need to be. RL is now typically only used to train reasoning for tasks with a well-defined correct answer (that's what I meant by binary reward) - this is called RLVR (RL with verifiable rewards).
Preference optimization (training the model on user "this response is better than that response" type data) is more often done with something in the same family as DPO (direct preference optimization), which is decidedly not RL.
Your philosophical concerns are correct of course. And there's the added caveat that the models that most people are using are closed, so we don't actually know their training recipes for sure.
A binary response vs rating is not related whether it learns confident or hedged tone. Either will produce a confident tone because humans respond more positively to a confident tone, hence the conman's language. Binary or not humans reward the tone and very much bias the model.
But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident. When the prior is greatly biased, a random number generator biased to that prior does better. The difference with humans and machines is humans are less likely to respond if they are less confident because they understand not knowing, which is why the training set is biased. It is one of the many fundamental flaw of LLM training and confusion of LLMs with intelligence. And that will not be fixed within the LLM architecture.
Yes, I am aware. I was referring to the fact that at this point, user preference optimization is not done with RL, but with other techniques. I was unintentionally being pedantic.
> But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident.
Yes, I think this has more to do with it than user preference optimization, honestly. "Valuable" text (i.e. text that produces a "good" model) for pretraining has the characteristic of being confident. Even models which are not optimized for chat (i.e. definitely no user preference data used to train them) exhibit this characteristic for medical questions (I know, because I've literally tested them for this purpose).
Preference optimization (or even RLVR) might play some small role as well, but it's kind of a "turtles all the way down" type problem.
Oh okay, darn, guess they just forgot to make it so!
Would it matter if this digital friend is not a real human behind a computer screen, but a Language Model in a data center?
I guess it falls into a similar category as buying "special performances to satisfy certain urges". It probably feels close to the real thing (I wouldn't know, I've never tried - promise! :P), but it's never the same as love.
GPUs are not people, and generated tokens can't have interest in a person's well-being. If you try to pretend otherwise, the results are not great. https://www.cnn.com/2025/11/06/us/openai-chatgpt-suicide-law...
I'm pretty sure you will be paid a large sum of money if you can make one of the frontier models urge you to commit suicide from normal interactions with it.
It's just a convenient way to ignore the years of evidence of the harms. Any bad news can be swept under the rug, labeled outdated as quickly as it happens. Well here's one that just happened, maybe this kid should have used a fancier model too? Should OpenAI pay him a large sum of money for the good he's done? https://www.cnn.com/2026/08/15/us/arjun-aravind-massachusett...
Reminds me of the 2000's crazy of "if you let kids play violent video games they will become unstable, violent psychopaths when they grow up". Thank God I was able to (rather easily) convince my mom that the idea that I would also steal a car and beat somebody to death with a golf club in real life just because that's what I did in GTA on my PlayStation, was completely absurd.
1. there weren't numerous real-life killings where the murderer credits GTA for coaching them through the crime
2. Rockstar Games didn't publicly say "sorry about the murders, but don't worry, we'll add more safeguards to GTA 6 to prevent even more people dying. GTA 6 will be the most aligned GTA ever!"
When a company admits to having blood on its hands, is maybe the point where things stop being "absurd" and start becoming real. But hey, what do I know. Maybe if HN existed in 2005 you'd have people commenting "who needs real people when we have computers, would it be so bad if lonely people have GTA as their only friend?" And I'd be the crazy one for engaging with them.
So. I think I agree with you.