No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration https://github.com/anthropics/claude-code/blob/main/plugins/...
and
>runs LLM inference repeatedly and saves output to markdown files for the next iteration
What exactly do you think self is in this case? What do you think the output contains? This history allows the LLM to have a more dialectic conversation with itself/agents to avoid iterating over the same problem space in a loop.
>anthropomorphizing
then please come up with a new dictionary for me to use that more accurately explains the behaviors exhibited in LLMs without just creating a parallel dictionary of different but equal words. I'll be glad to use it. No one has presented it so far.
Really?
> How many of our actions would look exactly the same.
Well, there are 8 billion of us, not exactly the same - a clear threat to humanity according to you.
> the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day.
This (statistically) doesn't happen with people because we follow rules. If you left your car on neutral and it struck another car in the parking lot, you broke the rules, not the car.
> LLMs are just an intelligence we can make a lot of copies of.
LLM's aren't intelligence, they are mechanical parrots of human knowledge. The meat parrots aren't all the same either and they don't repeat the same words for the same prompts, so what, ban parrots? If you left a bunch of parrots out of the cage, attached their beaks to gun triggers and they caused harm - you broke the rules, not the parrots.
LLMs are cool tech, AI can do amazing stuff, but if we let companies and people run amok and whenever something goes wrong we put the blame on these little independent angels with no accountability, we are one disaster away from a very tough spot.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
There a certain cargo cult of people just in denial about the impact/power of AIs.
Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
Now AIs are autonomously hacking websites, and people like to minimize the danger and blame it on the sysadmins.
I wonder what's going to be the next fad.
A stochastic parrot with human knowledge, using human tools, running human-designed trial-and-error experiments can solve human-defined hard math problems. This is nothing new, automated mechanical proofs predate LLMs by many years, you being unaware of it doesn't change it.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
Human failures all around, though it's easier to just blame the models.
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
Theatrics aside, anticipating zero days isn't only possible, it's required, even for the unknown ones. It wasn't that long ago when the AI labs were spending millions of $$ running their models to find multiple vulnerabilities, they even argued that they don't have to follow responsible disclosure, so proud of themselves in their privileged hubris.
At that time, no hacking happened because the models didn't have access to the wide internet, they were confined to a local computer or cluster.
In the HF case their unaccountable hubris went even further - the engineers knew the models can find zero-days and escape, nevertheless they ran the "experiment" on a system attached to the internet - the hacking is entirely the fault of human engineers and managers.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
My position, and the position of a large number in AI safety, is that you cannot build an intelligence that is both general and safe. The closer you get to generalized the more options the system has to do things that are wildly unsafe beyond human imagination.
This puts the AI labs in a serious bind while holding a bag filled with billions of dollars of debt.
Worse this puts governments in a multi-polar problem where even if the big public labs get shut down, black budget operations have a lot of free reign to make agentic digital weapons. Governments are not well known to take a lot of responsibility when their weapons cause damage unless they lose.