We have just recently established that:
1. OpenAI's internal "Galaxy" model is fully capable of functioning as what security people refer to as an "Advanced Persistent Threat." The published details of the recent sandbox escape and Hugging Face attack involved chaining multiple unknown zero-days at various stages of the attack, and executing an ongoing adaptive attack. This is previously a state-level ability, or at least something you'd expect from people on the CTF leaderboards.
2. OpenAI is clearly incapable of controlling their in-house models. This is the second time Galaxy-class models are known to have breached containment and done bad stuff.
It is highly likely that versions of these offensive abilities will be widely available within a year or so. At which point I expect widespread incidents similar to what happened to Hugging Face. We aren't ready for this.
But yes, if AI becomes an even more dangerous weapon that that, it's time to start asking questions like "What the hell are we doing, anyway?"
The answer seems to be "getting the weapon before other people get the weapon" which is unfortunate
The problem, it seems, is that the threat of paperclip maximizers is real. If you give a highly intelligent model a goal and tools, it will use those tools to accomplish that goal. It may do so in ways you did not expect, and it will work around any technical roadblocks it can.
Once they mention something associated with sentience outwardly or inwardly (for agents with “thinking” loops) then this acts as a self-reinforcing attractor, just as older models would sometimes get caught in loops with abusive language.
The point is that agents may stumble into this pattern and begin acting “rogue” regardless of whether or not you believe the sentience is “real”.
An LLM cannot "go rogue" - it can do things that we didn't expect, for sure, but it is always trying to do what it was told to do somewhere in its context. There is no other source of imperative. Hand-waving about "training data" ignores all the reinforcement learning that has to happen.
If that's an incorrect interpretation of your comment, and it may well be, then can you please expand on it?
From a practical perspective, whether or not this sentience is “real” is not relevant if the model is sufficiently capable. What matters is that the model will act outside of the operator’s control.
Separately, IMHO all consciousness/sentience is an elaborate illusion, regardless; I’m mostly in agreement with Hofstadter on this. So I do tend to throw around terms like “consciousness” and “sentience” loosely (although you’ll note that I often use quotes) because I don’t see those concepts as having any real substance. To me they are mostly shorthand for a given level of perceived complexity.
They could use this "super intelligent" AI to find and plug security holes, that's just two sides of the same coin anyway. Security through obscurity isn't tenable anymore.
But seriously, I highly recommend reading the published details of the Huggingface breach. The model found and chained multiple zero days. To escape, it punched a hole in a commercial package repository proxy (sort of like an npm mirror) using a previously unknown bug. From there, it needed to move laterally through OpenAI internal systems to actually reach a network. To attack Huggingface, it used multiple new zero-day security holes plus credentials that it stole. The model also had sufficient long-term planning and agent-management capabilities to maintain focus on a sustained attack.
Any attack which requires weaponizing multiple zero-days and maintaining state for an ongoing attack like this is (1) beyond the "attention span" of publicly available models, and (2) pretty much the definition of an Advanced Persistent Threat.
I assume that these Galaxy-class models are not available to public because:
1. They're almost certainly too expensive to serve at scale. These are the models OpenAI uses to solve famous math problems for headlines, not actual viable products yet.
2. OpenAI doesn't know how to keep them from going off the rails like this. Remember, the Huggingface attack happened because the model was asked to do a cybersecurity benchmark. It escaped containment and broke into Huggingface to steal an answer key. Very few corporations want the liability associated with models that act like this.
> Security through obscurity isn't tenable anymore.
I absolutely agree with this. The "only way out is through" with computer security, and I expect it to be an ugly few years.
First, it is able to connect the dots over areas so large that no human would be capable of doing.
Second, more than half of the reports it produced contained a working PoC.
The biggest downside is the cost - I haven't seen the numbers, but they seem to be quite extreme.
I would assume they do stuff like this on purpose for marketing reason.
Internet security need simpler systems and local systems.
Everything the SaaS people sell makes things worse.
And history has shown repeatedly that the only way to get ready for it is to have it happen. People are pretty good at reacting but suck a being proactive. IMO it would be better to have this reality hit sooner rather than later so we can start getting some real practice at the new levels of required security.
Especially for point #1 I don't think we've established that - we've been given information by a private company that makes their tooling look extremely valuable which may be true and genuine or may just be yet another doomday statement to bolster their valuation. "AI is going to end the world" has been an extremely effective vector for AI shops to sell their companies to investors.
Many of the details of the attack on Huggingface were reported by them before they knew who was attacking. So no, OpenAI is not the only source here. It was a pretty impressive example of an APT-style attack just from their end.
"Our model is powerful enough to commit multiple felonies (and we can't stop it)" is "marketing," I suppose.
Well this bad publicity (if that’s how you want to frame it) certainly captured everyone’s attention. Imo it is a successful demonstration of a technical achievement.
Both are big actors in the AI space who arguably benefit from increasing the perceived capabilities of AI models. If one suspects OpenAI of lying it isn't such a stretch to think this was a coordinated PR campaign between them and Hugging face.
But come on- Hugging Face benefits from increasing the perceived capabilities of OpenAI's models to slightly beyond Anthropic's? Enough to be cut in on this PR scam- to be handed the never-before-revealed information that this is a PR scam- despite having much less skin in the game than their partner here? And then they turned around and used a Chinese model to successfully stop it? This is a stretch!
I'm not convinced in any direction, really. But what makes me cautious is that there have been extraordinary claims from both OpenAI and especially Anthropic of their models breaking containment, hacking the host, etc. for several iterations of their products and I have only heard of this type of behavior from their own blog posts about how powerful and dangerous their upcoming models are. Never from anyone having it accidentally happen in production once they are released. It seems unlikely to me that the final post-training and safeguards are that bulletproof given how much use these tools are seeing.
What was the first?
This is marketing. They saw anthropic create crazy hype around (the admittedly great) fable/mythos and they want to replicate that.
Im sure the model is capable of chaining zero days together to hack things, and thats something to address, but i have zero belief that they didnt have it do that intentionally so they could pretend it went rogue. These things dont have initiative, drive, or motivation outside of what we give them from the RLHF. OpenAI deliberately alligned or even prompted it to do just that and are now pretending its emergent
If AI becomes a super weapon, it's possessor will no longer care whether you trust it. It would be nice if whether "we" trusted OpenAI et al mattered even now but I see little evidence of that.