But I wonder how much of that comes from RLHF itself or just from the way token prediction works.
And of course that is what it does, because there is no thinking involved! There is no logic. No consequence. No arithmetic. There is only continuation. An LLM can't continue a new idea, it can only continue a conversation about it.
An LLM does not have an opinion. Anything that looks like an opinion is just an emergent selection bias from its training corpus. LLMs are trained on what humans write, and human writing is kind and patient much more often than critical.
So what if we trained an LLM to be biased toward generating criticism? That would only replace the sycophant with a brick wall. What we really need is to find a way to bring logic and meaning into the system.