It’s based on users blind rating the output of two LLMs given the same prompt.
Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it.
And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just somehow subconsciously know how to write prompts that get me what I want.