The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.
I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.
(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf
Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.
We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?
I don't know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don't pay attention to them. That's not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.
When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don't care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.
Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.
Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.
Occams razor vs. Hanlon's razor?
OpenAI: shocked pikachu how could it hack?!?
They even went to a black hat conference and somehow boasted about it.
In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.
If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.
Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!
Truly psychopathic.
This is a very confident statement in the face of a purported non-0% chance of human extinction.
For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? https://www.theguardian.com/technology/2026/sep/28/ai-godfat...
I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.
There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.
At the same time, the technology is advancing in capabilities exponentially, and is beginning to exhibit long-predicted failure modes of RL that are nevertheless quite different than “insecure sandbox” or other that the software industry is used to.
The current crisis which OP claims is “fabricated” comes from the fact that the technology is advancing faster than anyone anticipated, the Hugging Face incident provides a clear example everyone can point to, and the labs have realized they cannot self-regulate because of a collective action problem.
There are a lot of levels of catastrophic damage that can happen between now and “long term hypothetical risks” like human extinction. When will it be worth regulation for you?
Do you use AI much? Not a trick question. What is your usage like? Reason I ask is recall OpenAI hyping GPT-2 with the same language as they are these new models. Doom and gloom. It sells. It gets attention. But use these systems enough and you start to become very familiar with their capabilities, and they are just so limited, and that's not even talking about how they lose the thread on long tasks. I understand that agency expands that a little, but not by much, really. When you use these systems a lot, it tends to be easier to see through the hype.
The real threat is, and will be for some time I think, people. Guard rails are not for AI. Guard rails are for people. All of this talk we are seeing in this space, this security chip being no exception, is treating a symptom, not a cause. People are what we should be focusing on, how, I have no idea, I couldn't even begin to guess how, but I can at least see that the actual issue is that the first thing some people want to do with it is cause harm and havoc. There is our sign. And we are trying to moderate the capabilities of the tool that can do harm in bad hands at a granular level when that means also moderating the same thing that can do good in that same tool. We aren't trying to moderate the hands that are handling the tool. And we need to be doing more of that. Any time we misapply constraints and focus on the shadow of the thing, and not the object casting the shadow, we are always going to be one step behind, whether it's AI or anything else. So I don't attribute the dangers of AI to AI itself, I attribute it to people.
Additionally, on AI building AI and AI just wiping us out; AI can do that for itself for rule-based things, up to god level I reckon. Things like Go, etc. And to be fair, AI research is partly like that, code either runs or it doesn't, so I get why that paper is worried. But where it needs data to actually succeed, it is limited by the amount of data that exists. And we, us, people, are who creates that data. AI and humans have more of a symbiotic relationship than people seem to realize. There are several papers that go into detail on this:
Models trained on their own output degrade: https://www.nature.com/articles/s41586-024-07566-y They need fresh real data every generation or they go downhill: https://arxiv.org/abs/2307.01850 Reusing the same data loses value after a few passes: https://arxiv.org/abs/2305.16264 And the stock of human-written text is finite: https://arxiv.org/abs/2211.04325
So it's not that I think it's 0%. I just think the explosion story underrates how much it still needs us, and that the damage in the near term is going to come from people, not the AI deciding on its own.
Indeed! They want the protections of our tax dollars because they have nothing else.
But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.
It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.
Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.