Reinforcement learning’s foundational flaw
thegradient.pub
thegradient.pub
“What are you doing?”, asked Minsky.
“I am training a randomly wired neural net to play Tic-Tac-Toe” Sussman replied.
“Why is the net wired randomly?”, asked Minsky.
“I do not want it to have any preconceptions of how to play”, Sussman said.
Minsky then shut his eyes.
“Why do you close your eyes?”, Sussman asked his teacher.
“So that the room will be empty.”
At that moment, Sussman was enlightened.
"picking the right reward function" in RL is shockingly hard. It actually works OK-ish when the problem space is strictly bounded, like with a game whose rules are known.
After that, you start getting into sky-humping cheetah problems:https://www.alexirpan.com/public/rl-hard/upsidedown_half_che...
https://www.alexirpan.com/2018/02/14/rl-hard.html is a better article, perhaps, than this one.
You can argue, but it works. Goal accomplished. Nature did things that are stranger than that.
Now you just have to add efficiency aspects into your reward functions and observe how it slowly finds a local minimum. Also something nature did a lot of times. Now hope that the remaining inefficiency is ok for you.
Done.
-
You can start with a random neural net. It's not exactly empty, but it's ok. The randomness defines which local minimum you'll find this time around while you burn through your VC money desperately hoping that the AWS bills for all the GPU time will arrive after you've found the holy grail that makes your startup worth <insert bullshit valuation>.
> "picking the right reward function" in RL is shockingly hard.
Agree. x1000. And computational power and data quality.
-
on another post I've added some links (AI playing Mario): https://news.ycombinator.com/item?id=17489459
One interesting thing about random NNs is how much they can already do. For example, you can do single-shot image inpainting or superresolution with an untrained randomly-wired CNN: https://dmitryulyanov.github.io/deep_image_prior In RL, you can use randomly sampled NNs to create various artificial arbitrary 'reward functions' to force an RL agent to explore an environment & learn the dynamics, and then when you give it the real reward function, it learns much faster how to optimize it. Similarly, you can sample random NNs to execute for entire trajectories for 'deep exploration', providing demonstrations of potentially long-range strategies much more efficiently than simple random-action strategies. In 'reservoir sampling', as I understand it, you don't even bother training the NN, you just randomly initialize it and train a simple model on the outputs, assuming that some of the random highly nonlinear relationships encoded into the NN will turn out to be useful, which sounds crazy but apparently works. Makes one think about Tegmark's interpretations.
in those situations you cannot expect RL to come up with “normal” solutions that do what you actually want. The sky-humping is a totally valid answer to the question asked; it’s merely that the reward function for “walk forward” doesn’t (and in many situations, may not ever) fully constrain the search space such that you get sane solutions.
https://news.ycombinator.com/item?id=16383264
Makes for an interesting trail of additional breadcrumbs to wind into the stack :)
No mention of all the ongoing work in learning from demonstrations, or more generally incorporating any off-policy knowledge. Vague speculations about the philosophy of model free learning. Not really worth the read (as someone working in RL).
Says as much at the end... to be fair we did warn up front "The first part, which you're reading right now, will set up what RL is and why it is fundamentally flawed. It will contain some explanation that can be skipped by AI practitioners." But personally I think the board game allegory is fun and that most people tend to forget the categorical simplicity of Go and Atari games and overhype ; easy to say the main points are not new but the details are important here.
Captioned chart:
The progression of AlphaGo Zero's skill. Note that it takes a whole day and thousands of lifetimes' worth of games to get to an ELO score of 0 (which even the weakest human can achieve easily).
I'm pretty sure that a one week old infant's ELO score will also fall short of 0. Sure, the AI did things that no human could do in order to match and then surpass human performance. Great! Half of the fun of following AI research is seeing it refute old intuitions about how human-like systems have to be to perform well on tasks previously considered to require human intelligence.
Whatever "general intelligence" or "human level intelligence" comes to mean by the 2050s, it looks like it's going to be a lot better pruned-by-counterexample than it was in the 1950s.
It's like 86 billion guys that try to please that thing they simultaneously produce (our consciousness). What I want to say: The algorithm can be dumb as f*. I call it the f-star-algorithm. But the computational power in our brains is extremely high.
So it's not only computational power, but also the unique structure nature found through trial and error.
But I did try to not just denigrate this work, rather to both praise it and discuss the sometimes ignores flaws.
> "Though DQN is great at games like Breakout, it is still not able to tackle relatively simple games like Montezuma's Revenge"
Yep:
https://www.engadget.com/2016/06/09/google-deepmind-ai-monte...
https://blog.openai.com/learning-montezumas-revenge-from-a-s...
It's far too early in this research to say what exactly what can and can't be solved by RL.
(Because the RL algorithm doesn't have access to the rules by which the simulation is carried out, it only has access to the commands and the result.)
And frankly, that would be a perfectly fair and interesting classification problem, so I don't see your point.
Otherwise, how exactly do you propose learning to drive a simulation without access to the simulation? I really don't know what you're saying here.
Thanks for your analogy though. I agree that it's better than mine. I was only trying to give a rough idea, but I'll use your analogy if I have to now. :)
In a sense it's pretty similar to how you'd learn a game if you watched someone play it through once. (Except backwards, perhaps.)
"Even 5 years later, no pure RL algorithms have cracked reasoning and memory games; on the contrary, approaches that have done well at them have either used instructions <link> or demonstrations <link> just as we mentioned would make sense to do in the board game allegory."
Newer approaches have the agent learn "primitives" through curiosity. It's sub-goal is to predict future states given the current state + an action.
By doing this, the problem becomes more hierarchical and the search space is reduced. This makes it feasible for more complex scenarios.
I haven't personally heard of a lot of research on this part but I imagine that transfer learning becomes more feasible as well once some "primitives" are established.
It seems like there's a bit of research in this area but it's not receiving the attention it may deserve. At least, that's how I interpreted the author's tone.
For background, here are some selected quotes from the article:
> "The first part, which you're reading right now, will set up what RL is and why it is fundamentally flawed." > "In the typical model of RL, the agent begins only with knowledge of which actions are possible; it knows nothing else about the world, and it's expected to learn the skill solely by interacting with the environment and receiving rewards after every action it takes." > "how reasonable is it to design AI models based on pure RL if pure RL makes so little intuitive sense?"
To summarize, the article claims that this particular aspect of RL is a "flaw".
I'd suggest it is more useful to call it a design choice. In many cases, this design choice has beneficial properties.
Of course, there are other ways to build learning agents. The field of RL is certainly open to alternatives, including hybrid models and/or relaxing this particular assumption.
I've seen a good number of (popular) articles about RL making rather broad claims, like this article. It appears to me that many of these articles attempt to 'reduce' RL to a smaller/narrower version of itself in order to make their claims. I hope more people start to see that RL is a set of techniques (not a monolith) that can be mixed and matched in many ways for particular applications.
Here is a quote from the article I want to mention: “Trying to learn the board game 'from scratch' without explanation was absurd, right?”
No. It is hardly absurd. Sometimes it works, sometimes not. It is a great starting point, if nothing else. So, I wonder if we have different ideas of what ‘absurd’ means.
I agree that we’re in a period of hype. It requires careful work to write clearly without too much zeal or oversimplification. My opinion here is that your attempt to ‘balance’ the debate uses a lot of language that I (and others) perceive as exaggerated.
Recent RL research about Policy Gradients / On Policy vs Off Policy / Function approximation / Model-based vs model-free are all research about how to get good at something with a lot of practice. RL has been around for a long time, discussions about higher level learning / planning has been done over and over. One doesn't discount the other. One deals with how to structure the learning problem that you can continue to get better with more experience (RL problem), while the other is about how to use higher level learning to speed it up.
"In part two, we’ll overview the different approaches within AI that can address those limitations (chiefly, meta-learning and zero-shot learning). And finally, we'll get to a survey of monumentally exciting work based on these approaches, and conclude with what that work implies for the future of RL and AI as a whole. "