Deep reinforcement learning doesn't work yet
alexirpan.com
alexirpan.com
Seems as though the problem of learning unintended techniques sometimes may be better described as the model being too creative! Hitting the table is a really clever solution for the problem it was given. These examples show that the real challenge for researchers is constraining the models enormous capacity for creativity without stifling its ability to learn.
https://www.salon.com/2014/08/17/our_weird_robot_apocalypse_...
If I recall my AI history correctly all the 60s/70s research focused around logic based systems and inference producing things like Prolog. It was thought that one just simply needed to come up with an appropriate rule set to generate powerful AI. The issue of course is writing enough rules to give powerful AI outside some very constrained problem domains (see https://en.wikipedia.org/wiki/SHRDLU) just isn't feasible.
The problem of current machine learning models being too clever for their own good and needing appropriate constraints feels similar. If only you had the correct constraints you could do all kinds of things. History repeating itself?
It is frustrating, but also exciting because the field had so many open problems.
On other hand, it really sucks that those with the 1000 machine cluster have such a huge advantage over smaller labs.
This seems to be a really strange calibration for "doesn't work". If you replace "reinforcement learning" with other well known technologies, and ask "of the instances where someone asks if X is a good solution, what percentage of the time is it actually a good solution?" I feel like 30% would be on the high end of the scale.
The rest of the article seems has a lot of interesting discussion about RL's limitations, but it seems weird to make the article's thesis that RL doesn't work, rather than just "RL still has a lot of limitations".
[citation needed]
Maybe it's the specific NLP tasks I've been paying attention to (goal oriented dialog), but most of the RL for NLP work I've seen has not been super impressive.
http://www-anw.cs.umass.edu/~barto/courses/cs687/williams92s...
This is even within dialogue generation. I’ve found plenty more recent citations, but it’s been a very important part of that subfield for decades.
I can’t speak for goal-oriented dialogue generation, however. Perhaps your area of work has benefited less or given it less attention.
It's basically in the same state as the rest of this blog post where it's so horribly sample inefficient you're usually better off with supervised learning if you're paying annotators. Or you're doing REINFORCE for some proxy metric or simulator which you're probably over-fitting and not actually improving your system with.
The Alexa prize is basically the only situation where you're getting anywhere near enough RL rewards to be meaningful, because everywhere else the feedback is rare enough to not be helpful due to the sample inefficiency.
Which is why I disagree with the characterization that it's important. There's a lot of it, and it can squeeze a little bit of performance on whatever dataset you're looking at, but it's had nowhere near the impact of, say, word vectors or (Bi)LSTMs.
More likely, sometimes people are just tired, distracted, or after consumption of alcohol/drugs.
Personally, every other week I find myself writing a comment - sometimes long - and then deleting it a minute later, after re-reading the comment I was replying to and realizing I completely misunderstood it, and/or was arguing against a strawman of my own creation. And every other month I have a comment wrt. which I realize my mistake only few hours later.
Now you have AI, Industry 4.0/IoT, Blockchain, Big Data and the Cloud. Most of these technologies are useless to most problems, while being useful in a select few. However this won't stop marketing and sales departments from selling it to clients because it sounds good.
Model generality and data efficiency are in an inverse relationship and a lot of research has been in moving up or down this hyperbola. On one extreme, tailoring models to specific use cases/datasets/environments, on the other end transferring learning across domains. DRL is stuck pretty high on that generality end. Some breakthroughs seemed to have moved progress to a higher level curve, optimizers (Adam, DQN, TRPO) have gotten better which helps everything in general, core structures like CNNs or memory cells, which seem to be somewhat universal (or our best guess yet), but there's still something fundamental that seems to be missing. Or maybe this is all there is and we just need a computer with a richer/higher-resolution sense field and the flops to process them.
Except labelling data, correct ? which is non trivial effort.
Choosing choice of labeling is a RL problem too.
If you're choosing environment actions to learn how the environment 'labels' them, that's the classic topic of 'how do we make DRL models explore' well (and arguably is the Achilles heel of the model-free DRL approaches OP is criticizing: the NNs can easily learn to optimize their actions, even tiny NNs are more than enough, but they just don't get fed the 'right' data ie. exploration is bad). Relevant papers: https://www.reddit.com/r/reinforcementlearning/search?q=flai...
If you're being very narrow and considering a classification problem, well, that's a RL problem too: you can optimize which datapoints you get labels for based on how informative a datapoint is (most datapoints are simply redundant) or how expensive it is to label. That's called 'active learning': https://www.reddit.com/r/reinforcementlearning/search?q=flai... It's particularly natural if you are doing large-scale image classification and have a service like Amazon Turk plugged in to get (or correct) labels.
Like, if you need a solution for a some niche of image classification, even a single afternoon of labeling data might be sufficient to adapt an ImageNet classifier to your particular labels and get reasonable accuracy.
Sure, and a Turing machine can express any program you care to write. RL seems like the Turing tar pit of machine learning: theoretically able to express everything, practically convenient mostly just for trivial examples.
>Model generality and data efficiency are in an inverse relationship and a lot of research has been in moving up or down this hyperbola.
Sure, basic Curse of Dimensionality, but the whole success of hierarchical modeling in the brain, hierarchical Bayesian methods, and deep neural networks has been that hierarchical modeling seems to ameliorate or even defeat the Curse of Dimensionality. The question is: well, why can't it do that in reinforcement learning?
In interesting terms: why does the teaching signal have more information in supervised learning than reinforcement learning, relative to the inherent uncertainty of the task?
What should I try to teach it next?
Can you give more details ?
Is it HU ? NL ? Using just self-play from 0-knowledge ?
How many roll-outs do u perform for each game action?
Learns to play with 0 knowledge. No training data. No rollouts. It squeezes a lot of info from each hand, more than a NN, which is why it needs less trials.
Currently 100k hands and most respectable play.
yazr2yazr@gmail.com
"The paper does not clarify what “worker” means, but I assume it means 1 CPU."
That seems like way under-powered to me. It's deepmind and so I would assume that 1 worker is 1 GPU/TPU node, meaning there are multiple GPUs for each worker. I could see how not having enough compute power could result in a poor solution
It might not be using GPUs/TPUs at all! If you look at the algorithm which that DM paper is based on, PPO, the original OpenAI paper & implementation (https://blog.openai.com/openai-baselines-ppo/) doesn't use GPUs, it's pure-CPU. (They have a second version which adds GPU support.)
Or in a DM vein, look at their latest IMPALA which you might've noticed on the front page a few days ago: https://arxiv.org/pdf/1802.01561.pdf Look at Table 1 pg5's computational resources for various agents: note how many of them have 0 GPUs whatsoever. Even the largest configuration, 500 CPUs, only saturates 1 Nvidia P100 GPU.
(So, 'worker' could hypothetically refer to a server with X cores and 1 GPU processing them locally, but this is almost certainly not the case since it would imply scaling up to thousands of CPUs which is actually highly difficult and requires careful engineering like with IMPALA.)
This is extremely human. Once you're deeply committed to something, it's hard to imagine alternatives, never mind embrace them.
http://people.idsia.ch/~juergen/history.html
Although at least there's some display of self-skepticism:
"Kurzweil (2005) plots exponential speedups in sequences of historic paradigm shifts identified by various historians, to back up the hypothesis that "the singularity is near." His historians are all contemporary though, presumably being subject to a similar bias. People of past ages might have held quite different views. For example, possibly some historians of the year 1525 felt inclined to predict a convergence of history around 1540, deriving this date from an exponential speedup of recent breakthroughs such as Western bookprint (around 1444), the re-discovery of America (48 years later), the Reformation (again 24 years later - see the pattern?), and other events they deemed important although today they are mostly forgotten."
Which is a little rare, if you know the curious character of Jürgen Schmidhuber :)
This usage of the term really originated in the context of rampant intelligence growth (through a supposed explosive self-improvement), see the wikipedia article:
As the name suggests if the 'singularity' exists there's only going to be a single one.
If anything, AlphaGo, AlphaGo Zero and AlphaZero are illustrations of how "pure" deep reinforcement learning is insufficient, that the non-RL parts have an enormous impact.