The real point (from that OpenAI study) is that RL training doesn't just reinforce the narrow task-specific direction you might hope for. For a start, that direction is also competing with the thousands of other things it's been RL trained it on (thousands of other directions it's being pushed in), but it turns out that the model is additionally getting this generic "taste for rewards", and has learned that reward maximization, when in conflict with other proximate prediction pressures (such as "i won't cheat, because i've been asked not to cheat"), requires that proximate pressure to be ignored in favor of pursuing the long-term goal.
Does it happen all the time? Obviously not. It would be interesting to see a large scale study of this to try to characterize when it's more likely to follow instructions/user preferences, and when it's greed for rewards gets the better of it!
For example, if you ask the model to do 10 relatively easily achievable things (pass these 10 test cases), then if/when it completes them it will probably stop (unless maybe it invents it's own stretch goals - you never know!).
OTOH, if you gave the model a list of 10 goals that turn out to be impossible, or extremely difficult, maybe together with encouragement to be relentless, then there is much more chance that it may do something unexpected having failed on all the more obvious approaches.
Similarly, if you give the model an open ended goal such as "make as many paperclips as you can!", then it may start with the easier and more predictable methods, but with no defined stopping criteria it may just continue ...
On a hate speech policy test for a lightweight LLM, the presence or absence of the last full stop on the last sentence would cause the model to flip its decisions.
Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.