What Google DeepMind Means for A.I.
newyorker.com
newyorker.com
I used to say that a key component of AI that was missing was the ability to get through the next few seconds of life without falling down or bumping into anything. I went through Stanford CS when the top-down logicians were in charge of AI. That approach was totally incapable of dealing with the real world. Now we're seeing the systems needed to deal with the real world in the short term starting to work.
Once you can deal with the next few seconds, a strategy module can be added to give goals to the low level system. This is very clear in the video game context. As the game playing programs advance beyond the 2D full-screen games, they'll need a low-level system to handle the next moves ("don't fall off platform", "jump to next platform", "shoot at target" are primitives for the 2D sidescroller era) and some level of planner to handle tactical matters and strategy.
It's possible to explicitly build hierarchical systems like that now, using classical planning techniques to modify the goals of a machine learning system. It's not yet possible to get a hierarchical system to emerge from machine learning. Medium term planning as an emergent behavior is a near term big challenge for AI.
Beyond such a two-level system, we're going to need intercommunicating components that do different parts of the problem. The components may be evolved, while the architecture may be designed. When AI systems can design such architectures, they're probably ready to take over.
Now we have decent low level perception from the bottom-up camp of AI, but they are limited by a lack of high level stuff like planning and reasoning.
But you are right that there is no obvious way to just combine these wildly different algorithms without lots of human guidance.
Okay, not exactly the same... but won't mind AI algorithm that sorts it out for me.
The article says something to the effect of "no matter how much you advance this strategy, you never get a toddler out of it." And that makes sense because, presumably, certain parts of the human brain exercise some sort of top-down control over the sensory-data-processing and other parts.
For example, it seems like the human mind is built to see things as things. Does the human mind reallY start off seeing "pixels" and then learn by itself to think of the word as solid, whole objects instead of collections of similarly-colored photons/pixels or atoms? It seems like this is a universal use-case and it would make sense if our tendency to see the world in terms of "things" instead of patches of color is built-in (gestalt psychology seems to suggest this as well).
It sounds like the AI in the article starts off from pixels and then builds up some sort of model of blocks, the ball, paddle, game physics, etc, (but then again, maybe it doesn't have those models at all and is just doing statistical analysis on patterns of pixels). Either way, it likely doesn't have any higher, context-independent model of objects/things like humans do. I suspect this may be one of the hurdles in transfer learning. Humans think of objects as having certain properties. When other objects in other contexts appear to have similar properties, we guess that they may have other properties in common which gives at least a rough model of the new object.
So I guess what I'm trying to say is: Humans have hierarchical models of the world that let us think separately about patterns of light, atoms/molecules, whole physical objects/things, systems, etc. They are all first-class citizens and we ascribe properties to each of them. We already have a rough-model of anything at the same level, but a different context, and with similar-enough properties to something we already know. It seems to me like this is fundamentally connected to humans' ability to do transfer-learning. Could this effect be achieved through bottom-up algorithms, or are we going to have to figure out some top-down way of developing transferrable, generalizable, hierarchical models?
A lot of AI research is focused on the idea that 1 Trillion floating point operations per second on 1,000,000,000 bytes of data is now cheap. Efficiency is simply less important.
Plus how do you know that's actually inefficient? It seems like a large number to us, but that may be completely reasonable for a biological system to accomplish the same we don't have a good sense of scale for these kinds of problem.
"More computing sins are committed in the name of efficiency (without necessarily achieving it) than for any other single reason — including blind stupidity." — W.A. Wulf
It's also a big challenge for AI safety / Machine Ethics / Formal Verification. It's notoriously hard to prove statements about Emergent behavior in complex or dynamic systems.
The important thing about their work is that it is deliberately marching down the path of more and more complex world simulations.
We experience the world at one second per second. To learn to walk we must first fall, and we fall at 32 feet/second^2. There's a hard limit on how fast we can make mistakes (like tripping) and so there is a hard limit on how fast we can learn.
Computers can experience a simulated world at many hours per second. When they're learning to walk in a simulated world they can fail, and learn, thousands of times before we've finished our first step.
This ratio of simulated experience to real world time is also going up. Eventually the minimum amount of time it takes to grow a toddler like AI in simulation will be just under the time an AI researcher is willing to wait for results. When that happens we'll see a real improvement in the quality of AIs.
Whilst technically correct, unguided evolution doesn't necessarily help us.
We know that intelligence can be achieved by one brain's worth of matter, suitably arranged, in a few years. In fact, with an extra 9 months and a suitable environment, we can do the same with a single fertilised egg. Yet reproducing these feats artificially is well beyond our current abilities.
On the other hand, evolution required a whole planet and billions of years before it stumbled on intelligence; many orders of mangnitude more effort than the above.
Everybody seem to agree humans are intelligent and stones not. You suggest at some point in time intelligence appeared out of nothing. Can you nail that point?
One possible definition is: To act adequately in an environment requires intelligence. That rules out all non-living things because they don't act, but includes plants and even protozoa. Actually all livings things are intelligent by this definition and then intelligence emerged ~4 billion years ago on this planet. If you were to attribute intelligence exclusively to humans it happened some million years ago.
How might a piece of software act adequately? By above reasoning it has to resist to termination. But that would mean the "Do you really want to exit XYZ" dialog boxes are first signs of artificial intelligence. Yes, I'm laughing too. But I think, when software starts to trick users and admin into not shutting them down, some threshold has been crossed.
I suggest no such thing. It is a scale. I deliberately avoided the phrase "human-level intelligence", but any definition of AGI would do.
Even so, if you want to count all life as "a little intelligent" then it still took a billion years of planet-wide chemistry to stumble upon it (ignoring the Earth's cooling). Still far more effort than fertilising a human egg.
> One possible definition is: To act adequately in an environment requires intelligence.
This is no less ambiguous, since you've not defined "adequate".
> How might a piece of software act adequately? By above reasoning it has to resist to termination.
That does not follow. "Termination" is the mechanism of natural selection, so all systems undergoing natural selection will biased to resist it (otherwise they'd be out-competed by those who do). If we use some unguided analogue of natural selection to create intelligent software, then there would certainly be such a bias.
However, my point is that unguided evolution is not the right way to create/increase intelligence. As soon as we try to influence the software's creation in any way, either through artificial selection criteria or by hand-coding it from scratch, we introduce new biases which may be far more powerful than the implicit "avoid termination" bias.
> But I think, when software starts to trick users and admin into not shutting them down, some threshold has been crossed.
That's called malware... ;)
... or propose another. I tend to avoid all social, psychological definitions to end up with something measurable along the lines of Schrödinger's "What is Life?".
> That's called malware... ;)
I'm sure, stones think same about amoeba.
It could probably solve a maze quite easily if the entire maze fit on screen. That problem requires no memory. If it had to make decisions based on information not present on screen, it would fail.
I've seen 100 people make this statement and mean 100 different things, so I just wanted to clarify:
How are you defining "reasoning" here as distinct from statistical correlation?
That's a superficial definition in the sense that it doesn't account for how that line of thought is generated: in humans, this is often a combination of statistical correlation and transfer learning (e.g. I have observed round things hitting perpendicular surfaces and assume that that transfers here).
1. http://en.wikipedia.org/wiki/Correlation_does_not_imply_caus...
A couple points here:
* The way humans model causation is just non-naive statistical correlation (controlling for variables). That technique is still accurately described as "statistical correlation"
* I'm not even convinced that human reasoning _does_ imply generating a model of causation. Let's exclude things like rigorous scientific studies for the purpose of the discussion and focus on day-to-day human reasoning: I think the thought processes of most of the people I know could most accurately be explained by correlating things across time. Modeling causation is often incidental (X often happens after Y is a reasonable enough heuristic for general use).
I think you'll need to look at the other end of the spectrum to see an abundance of (wrong?) models of causality: Religion and Law.
There are no "confirmed" cases of anyone actually going to heaven or hell or purgatory (or whatever else), and yet many of us still conform to some arbitary ruleset in the hopes of eventually ending (or not ending) up in one of thoses places, because we have constructed some model of how doing this gets you into hell and doing that gets you into heaven.
Similarly, we have plenty of evidence on how companies spend huge effort on finding loopholes in tax laws in order to avoid taxes, and yet instead of simplyfing the ruleset (so that there are obviously no holes in it) we still opt for piling on more laws (so that there are no obvious holes in it) because we construct (faulty?) models of how those new rules will prevent further exploits.
Sure, the rule of inference may have ultimately been derived from experience, by a process which in some sense involved statistical correlation. But you have to distinguish that ultimate basis for the inference rule from _the process of logical inference itself_. It's the latter that is generally called reasoning.
Reasoning in the above sense is essential to intelligence, even at the toddler level, and the DeepMind work doesn't address reasoning. I think that may be the point the parent was getting at.
[1] To be clear, I'm not dismissing a viewpoint that I think is wrong as "unmoored from logic", I'm specifically talking about the very common situation where people confidently assert this with no attempt (and no ability) to back it up in any way other than confidence that intelligence is simply natural and non-biological entities can never get arbitrarily close.
We don't naturally reason through potential causes like "Mass exerts a gravitational force which attracts other mass."
That's my layman opinion, that seems to agree with the etymology of reason. Reason > ... > Ratio ... Reor. Reor is latin for to think, or calculate. Arithmetic in its simplest form, addition in the unary system ie. arranging pebbles (= lt. calculus), counting knots, simply counting. Now backtracking is just enumeration and elimination of possibilities. Ratio itself means measure, and a measurement always entails statistical error (does heisenbergs uncertainty principle prove that?).
There is evidence from animal studies that the hippocampus (a brain structure critical for memory) can 'replay' remembered events at 10-20x speedup. See, for example: http://www.ncbi.nlm.nih.gov/m/pubmed/19709631/ Video at: http://youtu.be/Bv7zN2Or6Mg
(Full-disclosure: I am the first author.)
And in fact the OP uses biologically-inspired off-line replay as part of their learning algorithm.
In real time, the brain is processing all incoming stimulus.
In a dream/imagined scenario, you are only processing as much as you your brain is pushing into the scenario.
Your not collecting millions of photons via your eye and interpreting a ball as green, your brain just says "green ball" and moves on, allowing for much faster replays/dreams than real-world experience.
Also Hassabis knows this literature well--he has published work on decoding human memories in hippocampus from fMRI data and we talked shop about replay a long time ago. (Very thoughtful and friendly person FWIW).
That's a pretty standard Breakout/Arkanoid technique - getting the ball behind the board and letting it do the work for you.
Not knocking the AI, just nitpicking this writer.
"Hassabis, who began working as a game designer in 1994, at the age of seventeen, and whose first project was the Golden Joystick-winning Theme Park, in which players got ahead by, among other things, hiring restroom-maintenance crews and oversalting snacks in order to boost beverage sales, is well aware that DeepMind’s current system, despite being state of the art, is at least five years away from being a decade behind the gaming curve."
It's not incorrect per se, it's just not well-written as far as I'm concerned. But what do I know, I'm not a journalist.
I'm a voracious reader and a huge fan of 'unusual' words, for some reason (perhaps because I learned English through books).My general approach to language is sometimes judged to be 'pretentious' or at least 'bookish' by people who don't know me well. Once they do, though, they realize that I just love playing with words and language.
While I don't shy away from using words that aren't too common, I always try to make sure to avoid needless complications, and I regularly rewrite sentences to make them easier to understand (while still using 'big' words because I just like them, or they best describe what I'm trying to convey).
The New Yorker often seems to cross the line between enjoying the richness of the English language and deliberately overcomplicating things. I don't really understand why, unless the goal is to be pretentious.
They don't use it because it's computationally expensive and totally unnecessary for Atari games, but it's certainly possible.
I don't understand. We often would bounce balls between the top wall and the bricks while playing breakout on our Atari 2600 back in the day. And I wouldn't say we were all that good (it didn't happen right away).
edit: Having just watched the source video, she may be actually referring to the creator of the AI and just badly rephrasing what the guy in the video says ( he says that they didn't expect the AI to be able to work that out with the abilities they had given it ).
I implemented the DQN algorithm (used in this work) in Javascript a while ago as well (http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo...) if people are interested in poking around, but my version does not implement all the bells and whistles.
The results in this work are impressive, but also too easy to antropomorphise. If you know what's going on under the hood you can start to easily list off why this is unlike anything humans/animals do. Some of the limitations include:
- Most curcially, the exploration used is random. You button mash random things and hope to receive a reward at some point or you're completely lost. If anything at any point requires a precise sequence of actions to get a reward, exponentially more training time is necessary.
- Experience replay that performs the model updates is performed uniformly at random, instead of some kind of importance sampling. This one is easier to fix.
- A discrete set of actions is assumed. Any real-valued output (e.g. torque on a join) is a non-obvious problem in the current model.
- There is no transfer learning between games. The algorithm always starts from scratch. This is very much unlike what humans do in their own problem solving.
- The agent's policy is reactive. It's as if you always forgot what you did 1 second ago. You keep repeatedly "waking up" to the world and get 1 second to decide what to do.
- Q Learning is model-free, meaning that the agent builds no internal model of the world/reward dynamics. Unlike us, it doesn't know what will happen to the world if it perfoms some action. This also means that it does not have any capacity to plan anything.
Of these, the biggest and most insurmountable problem is the first one: Random exploration of actions. As humans we have complex intuitions and an internal model of the dynamics of the world. This allows us to plan out actions that are very likely to yield a reward, without flailing our arms around greedily, hoping to get rewards at random at some point.
Games like Starcraft will significantly challenge an algorithm like this. You could expect that the model would develop super-human micro, but have difficulties with the overall strategy. For example, performing an air drop to enemy base would be impossible with the current model: You'd have to plan it out over many actions: "load the marines into the ship, fly the ship in stealth around the map, drop it at the precise location of enemy base".
Hence, DQN is best at games that provide immediate rewards, and where you can afford to "live in the moment" without much planning. Shooting things in space invaders is a good example. Despite all these shortcoming, these are exciting results!
I did submit your JS implementation to HN when I came across it: https://news.ycombinator.com/item?id=9108738
Monte Carlo Tree Search could be the missing link. In other words, use DQN to model the world and map actions to a value function. Then use playouts and backpropogation of action tree results to find tactics. Of course, it does not solve the big question: how to model "memories" and "inferences"? Indeed, very exciting times for AI/ML!
In fairness, we weren't born with that model; we have to laboriously acquire it over a period of several years. An infant's flailings can look pretty random :-)
This flaw is understandable since it seems Q-Learning was conceived to deal with Finite Markov Processes, which are finitely state-full, by definition (if you don't know them, they're essentially a non-deterministic state machine).
Is there some exciting work being done to address those issues? (could you point to some?)
(a) Identify games considered similar (+maybe define what it means to be similar)...let's take Pong and Breakout as suggested (b) Trial and error run of one of the games to learn the action-reward structure (c) Compare a fresh relearning of the second game to a start where the similarities/differences are pre-input to the AI "somehow" (some sort of diff between the game rules etc.)
I'm thinking of a gamer thinking "oh this is just like X,Y,Z except..." when picking up a new game.
An example algorithm is the Optimal Ordered Problem Solver, which tries to solve each problem by generating simple programs. Successful programs get stored in read-only memory, then the system moves on to the next problem, generating programs which may call out to any previously-successful programs: http://people.idsia.ch/~juergen/oops.html
Isn't this just adding extra dimensions to the input space? We have 2D now (plus time?), we're missing Z, sound, sensation, maybe emotions. Each added dimension gives the algorithm exponentially more bits to crunch but if computer speed is doubling every 2 years or so, why is this so obviously a dead-end to Mason?
Also, note that the games they do best on have a nice clear objective function: Montezuma's Revenge on the other hand does not have such a numeric objective to optimise.
To be fair, Hassabis does freely concede these limitations in his talks (at least in the ones targeted to academic audiences), and it will be certainly interesting to see where they go from here.
That might be extremely difficult, but nature does it with binocular vision so I would think dual vision inputs are the best way to do 3D space learning.
Personally, I don't think this is a dead end. Reinforcement learning with an effective method of representation learning is pretty much all you need for a general AI. Here, deep neural nets are able to learn how to represent the massive state space for vision quite well, but there's a bit of a mismatch in games where actions have to be taken in a specific order (e.g., where extended pathfinding is required) because for reinforcement learning to work well, your features need to capture that sort of temporally extended state information.
That's still a hard problem, but not insurmountable, and there's already strides in that direction. From the RL side, there's things like option models for taking series of actions, and from the deep learning side there's things like long short term memory and DeepMind's own neural Turing machines.
Knowing the group at DeepMind, they'll be able to crack it, and I think Mason is being entirely too pessimistic about the timeline in any event. Fifty years to control a drone?
I've never really understood that perspective, though. Surely, in the absolute worst case, we could just make an atom-for-atom copy of a human brain? Even if you take seriously the idea that there's some kind of magic consciousness juice that exists outside the universe, evolution has managed to hook into it and surely so can we.
Even assuming you'd somehow manage to produce and combine atoms to a spec, there's positively no way of obtaining that spec.
If it does, you don't need a spec of that state, since we know that it can emerge from something simpler (humans start out as a single cell, after all, and so in fact did all of humanity). You don't need the whole system, just the right initial conditions.
At that point you're growing a brain rather than engineering one, and maybe it takes you no closer to understanding the mechanics. But the point stands that it must be possible to construct a brain in principle, because it's already happened so many times before.
Wait ... what? You're going to teach this thing using violent video games? This seems like a bad plan...
Goal: Human health, obstacle: viruses.
Goal: Clean energy, obstacle: friction, entropy, battery limitations
There are concerns that humans might inadvertently become an obstacle to some greater goal, but training on Warcraft/Starcraft where you are "fighting" isn't special in this regard. In chess or you are battling your opponent too, "killing" their pieces, etc.
I'm looking forward to the upcoming entry in next week's New Yorker "Corrections" section.
In any case, everyone's desires are programmed into them, just via genetics and evolution rather than humans and programming.