And we all know how good OpenAI is at containing models during training...
We have lots of examples now of their model doing what they say is impossible.
Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?