This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.
> The second seems to forget that jailbreaks are available for every model
Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.
This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.
You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:
"2. OpenAI’s harness and network security controls were unintentionally [...] bad"