I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.
(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf
Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.
We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?
I don't know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don't pay attention to them. That's not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.
When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don't care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.
Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.
Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.
Occams razor vs. Hanlon's razor?
OpenAI: shocked pikachu how could it hack?!?
They even went to a black hat conference and somehow boasted about it.
In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.