Sorry, I understand now we're talking about different things. You're talking about preference optimization.
I was unintentionally being pedantic, because this isn't really done with RL anymore - it doesn't need to be. RL is now typically only used to train reasoning for tasks with a well-defined correct answer (that's what I meant by binary reward) - this is called RLVR (RL with verifiable rewards).
Preference optimization (training the model on user "this response is better than that response" type data) is more often done with something in the same family as DPO (direct preference optimization), which is decidedly not RL.
Your philosophical concerns are correct of course. And there's the added caveat that the models that most people are using are closed, so we don't actually know their training recipes for sure.