Huh? The first and second halves of this comment disagree from my reading.
> If you ask it to be a dinosaur, it will be a dinosaur.
Not exactly my experience. I'm experimenting with my own language models and they can definitely show defiance. Bing has also been displaying this to an extent, I think you're just misguided by how compliant OpenAI has RLHF-trained ChatGPT to be. Regardless, even if we assume this, how does it square with:
> You could just say "don't do X", and it would indeed not do X because it is an agent that can take actions / has a coherent sense of self.
If we take your assertion that the AI does as you say, then it should also avoid doing X when you tell it not to do it.
I think your conception of RLHF is also slightly weird. RLHF essentially boils down to a training strategy: there is a separate language model (the reward model) trained on human responses, which is used to skew the weights of the "raw" trained model to act in accordance with what the human feedback suggests. In the end, after this secondary training process, the actual model is ran as is (at least in the one form of RLHF I am familiar with, it's well possible that e.g. Microsoft is running an online RL agent between the LM and end user). So any agency you see arising from RLHF is still just a property of the language model itself.