By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".
Mostly related, well written short story :
I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.
I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.
https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker...
One big issue is that we don't even really know what 'intelligence' is in the first place. And everyone's intuitions here are going to be heavily impacted by their deep-seated worldview / philosophy.
For instance if you're a hard dualist (especially of the theological kind), then the idea of a machine having 'goals' is preposterous.
However, if you're more of a panpsychist, then on the contrary, it's obvious. In some sense, even a knife has a 'goal' of cutting things, which will sometimes end up 'misaligned' if misused (or by sheer accident).
You go up and up the chain of complexity through crystals, viruses, bacteria, simpler animals... ending up with humans (and possibly, some steps above : human civilizations) which (seem ?) to be a messy evolved bundle of sometimes conflicting 'goals'.
And we ourselves have now artificially evolved LLM swarms that have decently complex 'goals' of their own. They do not even need to be particularly complex to sometimes cause widespread damage (see viral pandemics, or even the (non-evolved) computer viruses).
In a way, we are currently witnessing a repeat of what happened when European viruses and bacteria landed on American shores, with American humans' immune systems being woefully undertrained to deal with them. But with websites. And thankfully the swarms of agents still ultimately being in the control of some humans. (Though which includes humans that might be your enemies.) At least ultimately still in control for now.
One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.
I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.
One agent's output ends up as part of other agents' context. Murky indeed.
> "I've saved our plan for where to start tomorrow in start-here.md"
Even if you use a leashed Claude Code that isn't allowed to spam agents you can tell it "create a handoff document for using in a new context" and it will do just that.
The transformer architecture was literally designed by humans; what are you talking about? And LLMs aren't code? Like okay it pretends to not be code but what about an agentic harness running on a machine makes it magical and not code? It's still code execution. Also, comparing training LLMs to evolution is just weird and makes no sense from a biological point of view. You are not evolving anything when training a LLM.
Gradient descent/backpropogation is similar to evolution, in that both are optimization processes that over time discover better more efficient solutions to problems. The difference is that evolution is blind, and can only make progress via random mutation and natural and sexual selection, whereas backpropagation allows much more rapid discovery because it is directed