What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.
And I don't doubt that, not in the slightest. But I've seen exceptionally smart people in one field being dumber than a random kid from around the block in another.
This incident is clearly at least 2 failures that could've been easily avoided: failure to communicate, and failure to investigate the logs after letting the "most dangerous" roam free.
No, it doesn't require creating a mock internet with an alert as a side effect. Their own "most dangerous" model could have probably told them this happened if they supplied logs to it.
If I'm here to give them a good image I'm not doing very well at that.