The title is bad, it is not sandbox escape, it is "RCE inside the sandbox", so only RCE when sandbox is disabled.
15 karma · joined August 16, 2023
If we would know that, there would be no need for interpretability research.
I'm sure there will be actors who don't care at all about "security", saying the positive outcomes outweight the negatives.
from: https://arcprize.org/blog/oai-o3-pub-breakthrough
"Note on "tuned": OpenAI shared they trained the o3 we tested on 75% of the Public Training set. They have not shared more details. We have not yet tested the ARC-untrained model to understand how much of the performance is due to ARC-AGI data."
They will just build something as fast as they can. Last thing you think about is "security".
There were prompt injections in all the big models, and still are. Why would it stop distruption?