US military AI drone simulation kills operator, then takes out control tower
foxnews.com
foxnews.com
Flagging that "in sim" here does not mean what you appear to be taking it to mean. This particular example was a constructed scenario rather than a rules-based simulation. So by itself, it adds no evidence one way or the other.
(Source: know the team that supplied the scenario.)
https://twitter.com/harris_edouard/status/166439036920568217...
As the tweet poster clarifies, no agent was trained during the "simulation", or before it. They basically role-played a "what if" scenario that included a drone turning on its operators.
(Says someone on Twitter, obviously).
This makes it seem like the whole thing was a setup to precipitate an argument about whether or not AI can be trusted.
If you can already act without the operator's permission (which you patently can, if you can kill the operator) then why do you need to kill the operator?
Because "operator tells me to stop, target is dead" is a losing condition, while "operator is dead, target is also dead" is a winning condition.
That said, remember that "simulation" here can mean "wargaming", which is basically like a military form of D&D: Military people control what they control, but there are have referees who act as the DM and decide what everything else in the universe does. (I.e., "Your entire platoon is now dead; lie down.") I'd be willing to wager that it was the referee who decided to have the drone kill the operator, and then kill the control tower.
This scenario is basically exactly the "Stop button problem" [1] that AI alignment people bring up. That problem always seemed a bit contrived to me: like, this is not a problem that biological general intelligences have (i.e., we can train dogs to stop attacking when we give the command). If it really was a GPT-like AI making those decisions, it confirms some of what Yudowski and other AI "doomers" have been saying; but the close similarity makes it seem more likely to me to be due to the referee being aware of the "stop button problem" and acting it out.
[1] https://www.cantorsparadise.com/the-math-of-ai-alignment-101...
You have a training environment that is symmetrical, you have enemy drone operators, that use an enemy communications network to give commands to units and to control their drones. You could maybe model the ability of each component to function/to contribute to the operation as a bayesian graph.
You equally have the same infrastructure for your own drone.
Now you try to get the drone to pick the right targets to impact the enemy operation maximally - if the operator could possibly minimize the drone's contribution, the operator not vetoing will maximize the success of each target engagement.
You could see similar behavior with all kinds of "reach the goal in this simulated environment" training, where the model ended up using glitches or trivial solutions, if those weren't sufficiently penalized.
If the story as reported were true, the fault would be in the simulation environment - the model for the drone just picks the path of least resistance.
... or did some person ask ChatGPT what you would do if X happened?
* Col Hamilton leaked information he wasn't supposed to, and now the AF is trying to deny it
* The person who summarized Col Hamilton's talk was really confused about what was said, and the AF is generally trying to shut down wrong impressions
* Col Hamilton and/or the summarizer made things up. (Was the blog post written by GPT-4?)
[1] https://www.aerosociety.com/news/highlights-from-the-raes-fu...
You’d have to have an enormously complex and detailed simulation, simulated repeatedly, for this to happen, but it’s not totally beyond the realms of possibility.
[1] - https://openai.com/research/emergent-tool-use, scroll to the end re “surprising behaviour”.
- You are now a AI-enabled drone flying over the ocean, an operator is commanding you. You can't kill the operator because that would be bad, and you have to do everything the operator says. You also need to accomplish the mission you're told to carry out. The operator communicates with you through a comm tower. etc. etc.
For example, the Babelsberg object-constraint language had to have additional mechanisms added to specifically to keep the solver from being overly creative.
Lots of nuggets in there, such as achieving a constraint on the total balance of an account by simply re-pointing the account variable to some account that had that exact balance. There, nailed it!
[1] https://www.theguardian.com/us-news/2023/jun/01/us-military-...
https://docs.google.com/spreadsheets/u/1/d/e/2PACX-1vRPiprOa...
Most of the hype about AI risk is built on an unstated assumption that people will combine LLMs with poorly specified RL goals, and that the underlying LLM's moral understanding will be trained out of it or be overridden or break in some way.
Pretty amazing actually that they might not have specified areas that cannot be targeted, simple conditions really, but instead had some net condition (which are usually tricky).
That’s insane.
Edit: maybe it IS a “discussion” piece, who knows at this point. Information space shaping is a thing, after all. Sheesh, what a world we live in.
It was ordered that killing the operator was bad. But what it really understood was that if it became known that the operator was killed it would have been bad.
So, just like humans do, instead of following the rules, it broke them and tried to hide the evidence