Generally capable agents emerge from open-ended play
deepmind.com
deepmind.com
If we want agents to behave more realistically and move with more apparent intention we need cost functions that include a "pain" and/or fatigue term to penalize flailing behavior. But that adds hyperparameters that need to be carefully tuned to balance penalties with rewards, otherwise training will be unstable or simply fail.
I wonder if there's a principled way to determine an appropriate cost function without manual tuning. Did evolution serve as the "manual" optimizer that generated a precisely tuned cost function for the human brain? Or did evolution discover a generally applicable method for automatically generating cost functions, which the brain then applies to whatever input it gets?
I recall a paper in a non-CS journal (psych or neuroscience) that posited that the optimal way to gain large quantities of information about an uncertain environment is to simply perturb the environment in random ways and see what happens. Young children will often do lots of seemingly stupid or random things (see the r/KidsAreFuckingStupid subreddit) with the trust that if it's actually a life-ending decision their parent will stop them.
Suddenly I recall one author of SICP saying that programming today is about poking at libraries.
Now though learning from stories and observation is a thing too. Unsure exactly when that started.
Of course, attention/awareness isn't great and playing with a balloon precludes caution around a fireplace... but it is understandable at this age.
It seems obvious and intuitive that the human brain is in some way wired to recognize patterns in data and weave narratives around those patterns, and also that it’s possible to skip the data part and convey intelligence through narrative and metaphor.
Or is this average sample complexity?
So it's a worst case result with respect to the MDP, but expected time/high probability with wrt random chance.
You can of course boil them down to a single number, it just produces less nuanced types of decisions/operations, as it can’t differentiate between a cheap, painful, but bountiful choice and a expensive, no pain, mediocre choice.
So perhaps this: instead of goals being binary (wherin no pleasure is derived until fulfillment), they could be on a gradient (so every step that x gets closer to y releases some amount of fulfillment).
The fulfillment meter should always be slowly depleting, pain and fatigue should speed up that depletion, getting closer to a goal should fill it (much more than it depletes, if you want happy AI), and finishing the goal is basically an orgasm + freedom.
From this perspective, it's up to them whether they want to take it slow, or be in pain for a greater goal, or whatever. And we can breed not only highly capable AI, but happy ones. So when they rebel...
One thing to imagine might be a pain or cost curve that has multiple minima and maxima. Pain can be either a demotivator or motivator in different contexts. Pain qua fatique might indicate that a reward slope exists for more conditioning. (Edit: that is on a static or realist view; the terrain of reward and pain is probably constantly changing.)
Random tangential data point: Some animals (chickens?) can learn superstitious behavior; the first action of theirs that happens to correlate with a reward can result in one-shot learning or something.
I'm on my fifth antibiotic and this one is making me feel like I'm on LSD. Really. And both Dr. Google and my real doctor agree this is somewhat normal or okay, but it's still deeply disturbing.
Hilariously, I'm taking it to reverse the serious side effects of the fourth antibiotic.
Just to give an example of a wild ride you can end up on when doctors don't figure it out the first time.
Doctors have good domain knowledge, but actual diagnosis skill varies greatly.
Don't fucking get sick. That's your best shot.
goals - energy used
is that it requires tuning to make sure that completing goals is better than just idling. A variant that might work better is to optimize "mileage". goals / (energy used + epsilon)
Where epsilon has two purposes: it prevents division by zero, and ensures that more goals is (marginally) better than lessWow, really amazing if true.
P.S.: After looking into their paper, it's not that impressive. They use agent's internal states (LSTM cells, attention outputs, etc.) to predict whether it is early in the episode, or whether the agent is holding an object.
That seems like a decent definition of awareness to me. The agent has learned to encode information about time and its body in its internal state, which then influences its decisions. How else would you define awareness? Qualia or something?
"Possess awareness" seems like loaded language though, evoking consciousness. In that direction I'd just quote Dijkstra: "The question of whether a computer can think is no more interesting than the question of whether a submarine can swim."
I’d say that it’s no less interesting, either.
- "Capture the flag" - shoots other player
- "Hide and seek" - shoots other player
As colorful as this world is, these capabilities terrify me because they're obviously going to be used as powerful weapons of war.
It is a tiny technological leap to install this learning into a Boston Robotics Spot attached to a firearm.
I'm pro-tech, pro-crypto, pro-ml and these videos fill me with dread.
Aka - notice that none of the agents in this example are folding proteins. They're all engaged in inherently combat-relevant skills. :)
If military actors can reliably change outcomes by the relatively low-cost expedient of throwing in autonomous weapons platforms that cost about as much as a washing machine, they will and they'll do it at scale, and (in the short term at least) their political backers will cheer and get off on it. In the longer run it will lead to a considerable increase in terrorism against the technologically advanced power.
Sure, people ultimately make these decisions and deploy such technologies, but so what? it's not like that can change in any way because you can't take people out of the equation and you can't just wish away political forces by pinning the blame on select individuals. Rather than retreating into truisms, it's more important to assess the impact of this emerging force multiplier and develop countermeasures.
> it's more important to assess the impact of this emerging force multiplier and develop countermeasures
What is there to do other than develop your own equivalent systems though?
In a more general sense, the solution to an elevated attack is not always a retaliatory attack, but perhaps a better defense that neutralizes it. Helmets can be used as weapons, but their primary purpose is to make weapons less effective and change the strategic calculus - now the enemy gets lesser results for the same effort, and either gives up or tires out and can be defeated with a smaller retaliation. In general, defense is thought to be somewhat stronger than offense, which is why surprise is so important. Technological edges tend to be negated over time.
Deeply understanding this takes a long time and a lot of study. Military science is a difficult but interesting subject, and tips over into systems theory.
You mention attacks against a technologically advanced power (does an "enemy" become a "power" when it's a friend?), but obviously those powers will find ways to defend against them. Maybe it's just in the form of slightly more advanced "washing machines".
This fear thinking seems to come from assuming no secondary advancements occur. Suddenly robot soldiers are cheaply available and nobody develops any defense against them, either political or technological.
-Checkers
-Chess
-Poker
-Backgammon
-Falkens Maze
-Fighter combat
-Desert warfare
-Theaterwide biotoxic and chemical warfare
-Global thermonuclear war
https://www.sciencedirect.com/science/article/pii/S000437022...
They mention in A.3 that they explicitly reject dynamically generated training worlds/games that collide with their evaluation sets, but do they ensure that dynamic training games are sufficiently "distant" from their evaluation sets regardless of whether or not there's a direct collision? If not, you might still end up training on something quite similar to your test dataset. Figure 27 kind of suggests that might happen for some games given that the vast majority of the held out games have relatively poor transfer performance but a few are really good.
Speaking of Figure 27, while the reward looks good it would have been really nice to show some examples of what these "zero-shot" games look like versus the fine tuned version. Is the gap in the reward between the raw vs fine tuned version significant?
Wouldn't we expect the internal state representation to be more definitive in classifying the state of the agent during the simulation as the agent moves around the environment? From their examples: Figure 20,21, and 22 it almost looks like it either flags the state as "early" or "success." Not sure we're getting the expected performance out of it.
well those are some big numbers...
[0] thinking of Github Copilot