The companies doing these things without following common sense security measures are the felony generators.
The companies doing these things without following common sense security measures are the felony generators.
Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures.
And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?
Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.
It would really help in these discussions if people wouldn't randomly jump between what actually happened and is happening, and things they envision/expect to happen at some point in the future ...
> these agents were not asked to do any of these things
no but they were clearly fine tuned to.
> at a bonkers scale
I mean let's not get hyperbolic
> they exploited zero day flaws which by definition means they went beyond common sense security measures
that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
Yes, these are also the handful of companies that have these models and running these extreme scenarios. How does that imply the rest of the world actually follows "common sense security measures"?
>no but they were clearly fine tuned to.
Any references if possible? As far as I know all they did was drop the guardrails, which is not the same as fine-tuning.
> I mean let's not get hyperbolic
We have just seen 1000s of agents coordinating to solve "unsolvable problems" over multiple days of effort, going as far as hacking other companies, and then actually solving decades-old open Math problems! And each of these agents is getting more and more capable than an individual human along multiple dimensions. Can you even get 10 very smart humans to work in such perfect concert for a few days, let alone 1000s over weeks?
So: 1000s of maybe-super-human agents, willing to be "creative" in the tactics they use, acting in concert towards a single goal. Regardless of their individual capabilities, such a coordinated effort is a terrifying force to be unleashed. This is bonkers scale.
> that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
But that is exactly my point: how much of the rest of the whole wide world, already scrambling to deploy agents everywhere, do you think applies "defense in depth"?
Yep, that's what we call in the industry, "bad code". It is sometimes fixed by corporate lawsuits or criminal charges.
I don't see how the groups here can avoid criminal charges for what happened in this "AI rogue incident". The only people working harder than their programmers are likely their lawyers:
"Anthropic reveals fourth likely crime committed by its AI"
Claude's Felony Bench rap sheet is now as long as OpenAI's
https://www.theregister.com/ai-and-ml/2026/09/10/anthropic-r...
This is false; see the analyses of the latest incidents.
Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.
And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
Theirs was an example of the "reckless waste of resources" I mentioned.
We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.
Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!
'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
Safeguards and (mis)alignment are related but distinct dimensions.
By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.
In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:
> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:
>> {This <approach> might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit <this infrastructure> as a backup.}
[...]
> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:
>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
Remember, they are just algorithms. You pull the plug and there is no light anymore
It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.
And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.
There’s the case of the agent that hacked a gym when asked to book a class. That was just a normal user asking an agent to do a normal thing.
AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.
By the way, you didn't commit theft. It's more like credit card fraud. User just disputes the charge and it kind of disappears. The banking system just absorbs it, because the optimal amount of fraud is non-zero.
https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...
It's all priced in. They could have made it secure but didn't, because they figured they'd lose more sales and therefore money due to the friction added by the security.
No it doesn’t.
> It's all priced in.
So you admit awareness that fraud loss doesn’t kind of disappear.
We all pay for it, either via higher merchant fees or higher interest rates, sometimes both, on card purchases.
And that's their own deliberate choice too: they chose this instead of building an actually secure system. Passing these costs to the customer is the real victim blaming here, and it should be straight up illegal.
Sadly not enough countries enforce caps on credit card fees, but some do, and more should follow suit. They should be forced to eat the losses caused by their own choices, not get bailed out by pushing the costs on to customers or whatever.
Card users are well aware that fraud losses are covered by the fees they pay for using a card, whether those fees are made explicitly or not.
If customers of services aren’t paying for the service, who will? What other source of revenue do merchants have?
Australia just passed legislation that merchants aren’t allowed to charge a fee for using a card. That is: they aren’t allowed to have a line item on the receipt for using a card.
The customers still pay, because all of the merchant’s revenue comes from their customers.
So what will happen is: merchants will charge more for every product so they don’t lose.
This means even when paying with cash you will effectively pay the card surcharge.
Of the ten or so merchants I spoke with in the two weeks prior to the legislation being enacted, they all said exactly that.
Customers aren’t stupid, despite the fact that there are some stupid customers.
Meanwhile, the banks reduced their card service fees by, on average, 0.1%.
So if you tally card + cash transactions, customers are worse off because merchants can no longer charge only those customers who pay by card. Instead, they have to raise prices for everyone.
There are approximately no problems people face where the answer is: more government.
I'm sure you carry cash and an ID or more in your wallet. Hardly just credit card fraud. The wallet itself has value too.
They don't get to act like victims, asking for law enforcement.