An AI wolf that preferred suicide over eating sheep
lancengym.medium.com
lancengym.medium.com
Anyone who tries to devise optimal strategies for things should be able to see this isn't especially interesting.
Social metaphors are wildly out of place.
They say "unintended consequences of a blackbox" but I doubt that's true. Make it a deterministic turn based game and run it through a perfectly transparent optimization model and I wouldn't be surprised to learn this was just the best strategy for the rules they devised. I really hate when people describe an ai as something that cannot be understood because they personally don't understand it.
[1] https://inst.eecs.berkeley.edu/~cs188/fa18/assets/slides/lec...
It's interesting, though, how strong of a reaction general public had to this. The story must have strongly resonated with what some folks were already feeling. When you squint (pretend to understand the technology not at all) it's a tragic story. The situation of the wolf seems similar to the situation of some people. Chasing their careers in a highly structured, sort of dehumanized, environment of constant pursuit. "Supreme Intelligence" (that's what a layperson may think of AI) looks at a situation of the wolf and decides that it makes no sense to continue the pursuit. Moreover, what is "optimal" is the most tragic result - suicide.
Yes, because we don't see things as they are, we see them as we are.
Your perception IS your reality man!
Yes, yes, woe is the individual in modern capitalist society but the only reason people are reacting to this are that they don't understand it and they've been told it's something much more emotionally impactful than it actually is.
I think it's much more likely that they're reacting like this because they see their own plight in the wolf. It doesn't matter why the wolf killed itself, it became a meme that allowed many Chinese to reflect together on a common plight.
This seems like a common refrain when we see radicalized engineering students from less-developed countries, who are notably common in extremist groups. They're people on a very difficult path (an engineering program!) with no real path to success (living in a society where unemployment for people with degrees is very high). Cost for continuing on the path is high, and there's no obvious path to get the good outcomes.
> Perhaps the true lesson to be learnt here isn’t about helplessness and giving up. It’s about getting up, trying again and again, and staying with the story till the end.
I find the possibility of contrasting interpretations absurd. The problem with using any dead matter for our meaning making needs is it is ultimately a self-referential justification for how we think we should feel, while being equally or even more prone to self deception traps.
AI being the object is irrelevant here, this is nothing different than astrology or divination from tea leaves etc. It is 2000 BC level religious thinking with new toys.
The ONLY reason this was written was because the researches hired a programmer to make a specific thing, then is was too expensive for them to make more changes so they published the mistake.
This is a excellent analogy for this sort of behaviour, thank you.
I think that one thing it points to is how technology can discover novel iterations on a system. Imagine if this was a system modeled around a network and the agent was trying to figure out how to get from the outside to read a specific system asset. With the right (read: very detailed) modeling you could create a pentesting agent.
They have a word for it over there: involution i.e. no matter how much effort you put in, you get the same result.
Did my best to translate the (misguided) fitness function to fiction.
Cutting one's losses early may appear to be the most rational act if trying to minimize an agent's total suffering.
Factoid of the day for sure
Yes.
The survival of the actor has no intrinsic relevance to how the evolution develops.
No, not in this case. That was my point. That's why the outcome should not be surprising.
On the other hand, keep in mind that a significant weakness of most modern AI research is that it's extremely difficult to understand: you have the input, the output, and a bag of statistical weights. In the story, you know the (trivially bad) function that is being optimized; in general you may not. It's not without implications for other systems.
Further,
"At the end of the day, student and teacher concluded two things:
"* The initial bizarre wolf behavior was simply the result of ‘absolute and unfeeling rationality’ exhibited by AI systems.
"* It’s hard to predict what conditions matter and what doesn’t to a neural network."
> The initial bizarre wolf behavior was simply the result of ‘absolute and unfeeling rationality’ exhibited by AI systems.
This is a bad quote. They should not say this. It's a poorly trained agent doing a decent job of a poorly defined environment. Absolute rationality conjures images of some greater thinking but its actually a really stupid model that hit a local maxima. Calling it unfeeling implies the model has some concept of "wolf" and "suicide" but it does not. Replace the visuals with single pixel dots if you want an honest depiction of the room for feelings.
> It’s hard to predict what conditions matter and what doesn’t to a neural network."
This is generally true, but it isn't true here.
As part of my PhD research, I created a simplified Pac-Man style game where the agent would simply try to stay alive as long as possible whilst being chased by the 3 ghosts. The agent was un-motivated and understood nothing about the goal, but was optimising for maximising its observable control over the world (avoiding death is a natural outcome of this).
I spent sometime trying to debug a behaviour where the agent would simply move left and right at the start of each run, waiting for the ghosts to close in. At the last minute it would run away, but always with a ghost in the cell right behind it.
Eventually, I realised this was an outcome of what it was optimising for. When ghosts reached cross-roads in the world they would got left or right randomly (if both were same distance to catching the agent). This randomness reduced the agent's control over the world, so was undesirable. Bringing a ghost in close made that ghost's behaviour completely predictable.
I naively thought it would be some kind of Kalman filtering of sorts but from what I gather in your words it doesn't even have to be "that" complicated, right?
edit: found your link to the paper in another post ( https://news.ycombinator.com/item?id=27749619 ), thanks!
To achieve the Channel Capacity you need to find the optimum distribution across a - i.e. what set of signals maximises the information you can transmit on this channel. There are known algorithms for finding this distribution (e.g. Blahut-Arimoto).
Now if you model the world as a channel, where s represents the reachable states and a represents the actions the agent can take (and the channel, P(s|a), represents the dynamics of the world), you can calculate what actions allow you maximal control (in terms of states you can controllably reach).
More info in this paper: https://uhra.herts.ac.uk/handle/2299/15376
From a mathematical perspective, we used Information Theory to model the world as an information theoretic 'loop'. The agent could 'send' a signal to the world by performing an action, which would change the state of the world; the state of the world was what the agent 'received'. This obviously relies on having a model of the world and what your actions will do, but doesn't burden the model with other biases.
Pore more colloquially, the agent could perform actions in the world, and see the resulting state of the world (in my case, that was the location of the agent and of the ghosts). Part of the principle was that changes you cannot observe are not useful to you.
Having all the ghosts behind you gives you more control since they'll follow you in a line.
That the ghosts follow the player is what makes the game winnable. If they formed a grid and gradually closed-in, it would be impossible to escape.
Edit: What was unexpected in this case was that the system found a strategy the programmer didn't think of.
For example in EVE Online with a 1v1 fight two basic tactics are either Kite or Brawl. A kiter that can maintain range will beat a brawler. But a brawler that 'catches' a kiter will generally win.
It is a fun feeling when your own program surprises you.
The number of outcomes in branching_factor^t (very large) makes the action-values at t=0 (where the agent chooses between two/three actions) almost uniform random.
I experimented with different time horizons, mostly look 3-7 steps ahead.
In terms of the 'reward', that was implicit within the model - if the ghosts caught you, your ability to influence the state of the world dropped to 0.
The cars' "fitness" function rewarded cars for driving along the course and punished them for crashing into walls. But evidently this function punished a little too severely: the most successful cars would just drive in tight circles and never make progress on the course. But they were sure to avoid walls. :)
Edit: don’t want to sound accusatory
You can see the paper OP wrote to confirm for yourself that their story is not the same as yours: https://uhra.herts.ac.uk/bitstream/handle/2299/15376/906989....
That is very interesting that this emerged from two different approaches.
I published my result years back, and have never heard of this emerging elsewhere before!
Didn’t take it as accusatory [but thanks to child for sharing link :)].
https://www.lesswrong.com/posts/4ARaTpNX62uaL86j6/the-hidden...
In it, there is a thought experiment of having an "Outcome Pump", a device that makes your wishes come true without violating laws of physics (not counting the unspecified internals of the device), by essentially running an optimization algorithm on possible futures.
As the essay concludes, it's the type of genie for which no wish is safe.
The way this relates to AI is by highlighting that even ideas most obvious to all of us, like "get my mother out of that burning building!", or "I want these virtual wolves to get better at eating these virtual sheep", carry incredible amount of complexity curried up in them - they're all expressed in context of our shared value system, patterns of thinking, models of the world. When we try to teach machines to do things for us, all that curried up context gets lost in translation.
> Suppose we have an AI whose only goal is to make as many paper clips as possible. The AI will realize quickly that it would be much better if there were no humans because humans might decide to switch it off. Because if humans do so, there would be fewer paper clips. Also, human bodies contain a lot of atoms that could be made into paper clips. The future that the AI would be trying to gear towards would be one in which there were a lot of paper clips but no humans.
[1] https://en.m.wikipedia.org/wiki/Instrumental_convergence
Genetic Algorithms attempt to use this same system over extremely simple "fitness landscapes," where the fitness of an agent is defined by programmers using some simple mathematical formula or something.
When the fitness function is being defined in the system by programmers, instead of emerging from a rich and complex ecosystem, then the outcome depends exactly on what the programers choose. If they fail to see the consequences of their scoring algorithm, that's on them. There's nothing really magical going on, they simply failed to foresee the consequences of their choice.
(As someone who has worked with GAs and agent models, this outcome really doesn't surprise me. I would have said "oops, I need to weight the time less" and re-run it, and not thought twice.)
https://www.bilibili.com/video/BV16X4y1V7Yu?p=1&share_medium...
AIs need to learn to feel awkward and avoid it, just like we humans do (even if it feels very irrational at times).
https://www.nature.com/articles/s41599-020-0494-4
What he means is that computers, which can learn rules and use those rules to make predictions in certain domains, nevertheless cannot exercise general intelligence because they are not "in the world". This renders them unable to experience and parse culture, most of which is tacit in real time, and sustained by enduring mental models which we experience as "expectations" that we navigate with our emotions and senses.
Culture is the platform on which intelligence is manifest, because the usefulness of knowledge is not absolute - it is contextual and social.
But as for being a dualist in the 21st century, there is always consciousness, information and math. All three of which can lead to some form of dualism/platonism.
1. What is special about a body that makes it impossible to have intelligence without it? (a) Is it possible for a quadriplegic person to be intelligent? (b) A blind and deaf person? ((c)What about that guy from Johnny Got His Gun?)
2. What is special about a childhood such that a machine cannot have it?
3. Would a person transplanted into a completely alien culture not be intelligent?
What is fundamentally being argued is the definition of "intelligence", and there are many fixed points of those arguments. Unfortunately, most of them (such as those that answer "no", "probably not", and "definitely not" to 1a, 1b, and 1c) don't really satisfy the intuitive meaning of "intelligence". That, and the general tone of the arguments, seem to imply the only acceptable meaning is dualism.
For example, "...there is always consciousness, information and math...": without a tight, and very technical, definition of consciousness, that seems to be assuming the conclusion. With a tight, and very technical, definition of consciousness, what is the problem with a machine demonstrating it?
Information? Check out knowledge, "justified true belief", and the Gettier problem (https://courses.physics.illinois.edu/phys419/sp2019/Gettier....).
Math? Me, I'm a formalist. It's all a game that we've made up the rules to.
To me it sounds dualist if intelligence is disembodied. If the substrate doesn't matter, only the functionality, then that sounds like there's something additional to the world than just the physical constintuents. But of course, embodied versions of intelligence need to answer the sort of questions you posed. It should be noticed that Dreyfuss wrote his objections in the 50s and 60s during the period of classical AI. I don't know whether he addressed the question of robot children, or simulated childhoods. We don't have the sort of thing even today, and we also don't have AGI. Some of his objections still stand, although machine learning and robotics research has made inroads.
> Math? Me, I'm a formalist. It's all a game that we've made up the rules to.
So why is physics so heavily reliant on mathematics? Quite a few physicists think the world has a mathematical structure.
> For example, "...there is always consciousness, information and math...": without a tight, and very technical, definition of consciousness, that seems to be assuming the conclusion.
Qualia would be the philosophical term for subjective experiences of color, sound, pain, etc. Reducing those to their material correlations has been notoriously difficult, and there is still no agreement on what that entails.
As for information, some scientists have been exploring the idea that chemical space leads to the emergence of information as an additional thing to physics which needs to be incorporated into our scientific understanding of the world. That we can't really explain biology without it.
"To me it sounds dualist if intelligence is disembodied. If the substrate doesn't matter, only the functionality, then that sounds like there's something additional to the world than just the physical constintuents."
Off the top of my head, what the substrate is doesn't matter, but that there is a substrate does. Intelligence is the behavior of the physical constituents.
"So why is physics so heavily reliant on mathematics? Quite a few physicists think the world has a mathematical structure."
Because humans are very good at defining the rules when we need them? Because alternate rules are nothing but a curiosity even to mathematicians unless there is a use---such as a physical process---for them?
One of the problems with qualia, as a topic of discussion, is that I can never be entirely sure that you have it. I can assume you do, and rocks don't, but that is about as far as I can get.
If you put a computer in a room with a hot babe, a 3 layer chocolate cake, a bottle of the finest whisky or bourbon, the keys to a Porsche, and a trillion dollars in cash, what would it do?
Yeah, nothing. The computer is not in the world.
Assuming you are human, that depends on how long you care not about food or drink.
Yes of course, because all of those people have ambitions and desires. They feel pain and they seek pleasure, which they experience through their bodies.
Imagine if the world 2,000 years from now was populated only by supercomputers, all the lifeforms having perished.
What are these computers going to do with the planet?
How to have fun and avoid disaster? That’s a definition of intelligence.
Implies dualism. In a materialist world a computer can learn anything given the proper structure and stimuli.
Consider this line from an Eagles song:
“City girls just seem to find out early, how to open doors with just a smile.”
What does that mean to you?
Disembodied computers don’t get the experiences required to gain that intelligence, and even if they could go along for the ride, in a helmet cam, they wouldn't experience the tingling in their heart, lungs and genitals that provide the signals for learning.
Similarly, nuclear submarines, which lacking all of the critical organs of fish, are completely unable to swim.
I feel like that's utter nonsense. What the things we misname 'AIs' today don't lack intelligence. They lack motivation. Goals. It has nothing to do with childhood or culture. It's not What™ or How™ that is missing but Why™.
For example even the dumbest 'living' organism is motivated to reproduce. Even if it doesn't know why. But since all the ones that weren't didn't, they died out and all we're left with are the ones that do.
And humans without a Why™ strongly resemble what we call depressed.
The technical details aren’t interesting, but I do think it’s interesting just how disjointed life is vs what was promised.
In the US, this was aptly named a rat-race; and the white collar Chinese with a market-based economy are suffering the same.
Our markets and nations promise some combination of wealth or retirement and enjoyment of life, but it’s an ever-moving goal just out of reach for anyone but the lucky few.
(Obviously closed, specific tasks like "land this particular rocket safely within 15 minutes" don't always lead to this, but open ended ones like "manufacture mcguffins" or "bring about world peace" sure seem to.)
This one becomes especially dangerous after the 15 minutes have passed and it begins to concentrate all its attention on the paranoid scenarios where its timekeeping is wrong and 15 minutes haven't actually passed.
Which pretty much tackles these issues head on.
There are interesting videos on the subject on Robert Miles channel on AI safety: https://www.youtube.com/channel/UCLB7AzTwc6VFZrBsO2ucBMg
From there it's a sequence of steps that would show up in a thorough root cause analysis ("humanity, the postmortem") where the agent capitalizes on existing abilities to gain more abilities until murder is available to it. It would likely start small with things like noticing the effects of stress or tiredness or confusion on human opponents and seeking to exploit those advantages by predicting or causing them, requiring more access to the real world not entirely represented by a chess board.
What they are showing is one of the main issues with agent-based-models (and I think every model, but it happens particularly with models trying to capture the behaviour of complex open systems): Garbage in -> Garbage Out.
Most likely the representation of the sheep/wolf system was not correct (so the modeling was not correct). Here "correct" means good enough to demonstrate whatever emerging behaviour they are studying. ABM is a powerful tool, but you must know how to use it.
It’s easy to sort out in narrowly specified areas, but an extremely hard problem as the tasks become more general.
AI is worth calling out in this regard because, if the field is successful enough, it can create dangerous systems that don’t behave how we want.
Building a safe general AI is much harder than building a general AI, which is why it’s worth considering AI as it’s own problem domain.
If you add your penalty, and a deficit of nearby sheep, you'd expect a trifurcation of strategy: hoarders that consume the nearby sheep immediately, explorers that bet on sheep further afield, and suicides from those that have evaluated the -100 penalty to still be optimal.
Let's say you are a human player playing the wolf and sheep game. The score achieved in the game decides your death in real life. Note the stark difference. Dying in the game is not the same thing as dying in real life.
If there is an optimal strategy in the game that involves dying in the game you are going to follow it regardless of whether you are a human or an AI. By adding an artificial penalty to death you haven't changed the behavior of the AI, you have changed the optimal strategy.
The human player and the AI player will both do the optimal strategy to keep themselves alive. For the AI "staying alive" doesn't mean staying alive in the game, it means staying alive in the simulation. Thus even a death fearing AI would follow the suicide strategy if that is the optimal strategy.
It is impossible conclude from the experiment whether the AI doesn't fear death and thus willingly commits suicide or whether it fears death so much that it follows an optimal strategy that involves suicide.
For example, there's a small child who is learning to walk. The child falls down a lot. Eventually the child will work out a long list of arbitrary negatives connected to its wellbeing that are associated with falling down.
However, the parents, being impatient, reach inside the child's head and directly tweak some variables so that the child has more dread of falling over than they do of walking. Did the child learn this, or was it told ?
We currently do the latter every time an agent gets something wrong. Left to their own devices, 99.9% of agents will continue to fall down over and over again until the end of time.
We have a long way to go before we can say we've created 'AI'.
Is it artificial? Does it make decisions? It's an AI. Even if it's crappy, and not very intelligent.
I had my own expeirience with this when I tried to train "rat" to get out of the maze. I rewarded rats for exiting but for some simple labirynths I generated for testing it was possible to exit it by just going straight ahead. So this strategy quickly dominated my testing population.
One particular solution stood out: https://codegolf.stackexchange.com/a/25357
The suicidal wolf became a (short-lived) running gag so it started appearing in other king-of-the-hill challenges: https://codegolf.stackexchange.com/a/34856
One of these days I have to actually scour the web and collect a few good examples where evolutionary methods are used effectively on problems that actually benefit from them, assuming I can find them. Almost every example you're likely to see is either a) solved much more effectively by a more traditional approach like normal gradient descent or classic control theory techniques (most physical control experiments fall into this category), b) poorly implemented because of crappy reward setup, c) fully mutation-driven and hence missing what is actually good about evolution above and beyond gradient descent (crossover), or d) using such a trivial genotype to phenotype mapping that you could never hope to see any benefit from evolutionary methods beyond what gradient descent would give you (if the genome is a bunch of neural network weights, you're definitely in this category).
https://docs.google.com/spreadsheets/u/1/d/e/2PACX-1vRPiprOa...
From https://deepmindsafetyresearch.medium.com/specification-gami...
The glib answer is never, of course. And one easy-out, I can think of is setting a fixed/limited lifespan for the AI and maybe allow suicide or an off-button. So the AI can ultimately choose to 'opt-out' should it like; and at least, suffering isn't infinite or unending.
It reminds me of reactions to testing the stability of Boston Dynamic's early pack animal. The people giving the demo were basically kicking it, while the machine struggled to maintain its balance. The machine didn't have the capacity to care, but to a person viewing it, it looked exactly like an animal in distress.
Dismissing “never” offhand without explanation is glib.
Reward hacking is dangerous because the global optima turns out to be different from what you wanted, and the smarter and faster and better your agent, the worse it becomes because it gets better and better at reaching the wrong policy. It can't be fixed by minor tweaks like training longer, because that just makes it even more dangerous! That's why reward hacking is a big issue in AI safety: it is a fundamental flaw in the agent, which is easy to make unawares, and which will with dumb or slow agents not manifest itself, but the more powerful the agent, the more likely the flaw is to surface and also the more dangerous the consequences become.
Clearly in the (flawed) objective there is a phase transition near the very beginning, where the wolves have to chose whether to minimize the time penalty or maximize the score. With enough "temperature" and time perhaps they could transition to the other minimum, but the time penalty minimum is much closer to the initial conditions, so you know ab initio that it will be a problem. You can reduce that by making the time penalty much smaller than the sheep score and adding it only much later. I feel bad that the students wasted so much time on a badly formulated problem.
Edit: Also none of these problems are black boxes if you understand optimization. Knowing what is going on inside a very deep neural network (such as an AGI might have) is quite different than understanding the incentives created by a particular objective function.
"In the initial iterations, the wolves were unable to catch the sheep most of the time, leading to heavy time penalties. It then decided that, ‘logically speaking’, if at the start of the game it was close enough to the boulders, an immediate suicide would earn it less point deductions then if it had spent time trying to catch the sheep."
It's as if the scenario you are thinking about involves "assume a machine capable of greater-than-human-level perception, planning, and action" and then set it to optimize a trivially bad function.
How many people do you know with a single goal of "die with as much money as possible", which has a trivial solution: rob a bank and then commit suicide.
Manager to boss: It’s a crazy new AI behaviour that is going viral around the world!
This does not only concern AI systems, but all systems in general - including human ones.
The incentive structure is a two dimensional membrane embeded in a third dimension of "points space."
Obviously if the goal is to maximize total points OR minimize point loss and the absolute value of the gradient toward a mininum loss is greater than the abs gradient toward a maximum gain then the algorithm may prefer the minimum until or if it is selected against by random chance or survivorship bias.
obviously the linear time constraint causes this. a less monotonic, i.e. random, time constraint may have been interesting.
Here is the full video also linked at the bottom. It also shows the one that trained longer that the wolves start successfully hunting the sheep after more training examples.
Another interesting observation is that the wolves don't coordinate it seems. That probably implies that the reward functions are individual, so they're technically competing rather than cooperating.
Lastly... they seem to not be very good at the game even at the end
"Drawn by the fascination of the horror of pain and, from within, impelled by that habit of cooperation, that desire for unanimity and atonement, which their conditioning had so ineradicably implanted in them, they began to mime the frenzy of his gestures, striking at one another as the Savage struck at his own rebellious flesh, or at that plump incarnation of turpitude writhing in the heather at his feet."
We have so many systems in the real world that set up bad incentives for humans, yet the concept is largely misunderstood by politicians and decision makers. Our democratic discourse is dominated by first-order thinking, our laws are too often written under the assumption that the affected entities' behaviour will remain the same under the new incentives, which never holds.
I'm not buying that. As soon as they mentioned the 0.1 point deduction every second it seemed obvious?
https://news.ycombinator.com/item?id=5397797
Be sure to read to the end.
One of the answers is also pure gold in context:
> Don't feel bad, you just fell into one of the common traps for first-timers in strong AI/ML.
I'll pass thanks.
Or is that completely thrown out the window if you use a ML model rather than a procedural algorithm?
Because if the model is a black box and you use it for some safety system in the real world, how do you know there isn’t some wierd combination of inputs that causes the model to exhibit bizzare behaviour?
Here's a collection of such stories:
> " William Punch collaborated with physicists, applying digital evolution to find lower energy configurations of carbon. The physicists had a well-vetted energy model for between-carbon forces, which supplied the fitness function for evolutionary search. The motivation was to find a novel low-energy buckyball-like structure. While the algorithm produced very low energy results, the physicists were irritated because the algorithm had found a superposition of all the carbon atoms onto the same point in space. “Why did your genetic algorithm violate the laws of physics?” they asked. “Why did your physics model not catch that edge condition?” was the team’s response. The physicists patched the model to prevent superposition and evolution was performed on the improved model. The result was qualitatively similar: great low energy results that violated another physical law, revealing another edge case in the simulator. At that point, the physicists ceased the collaboration."
The problem was discovered when they couldn’t get the same results on a different FPGA, or in the same one in different day (subtle variations of voltage from mains and the voltage regulators).
They had to redo the experiment using simulated FPGAs as a fitness filter.
Maybe there's a repository somewhere with similar examples?
[1](https://towardsdatascience.com/today-im-going-to-talk-about-...)
Tho it is interesting how people in China related the broken rules of the game (that lead the ai to commit suicide) to the broken rules of their lives in a crushingly oppressive authoritarian nation.
Would be more realistic because dying has higher cost than failing.
Slap the term AI on anything and get automatic press coverage?