> A binary response vs rating is not related whether it learns confident or hedged tone.
Yes, I am aware. I was referring to the fact that at this point, user preference optimization is not done with RL, but with other techniques. I was unintentionally being pedantic.
> But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident.
Yes, I think this has more to do with it than user preference optimization, honestly. "Valuable" text (i.e. text that produces a "good" model) for pretraining has the characteristic of being confident. Even models which are not optimized for chat (i.e. definitely no user preference data used to train them) exhibit this characteristic for medical questions (I know, because I've literally tested them for this purpose).
Preference optimization (or even RLVR) might play some small role as well, but it's kind of a "turtles all the way down" type problem.