A JavaScript deep learning and reinforcement learning library
github.com
github.com
Wouldn't the libraries/algos be slightly different throwing off the weights?
sending the input data to the server, doing the computations there and getting the answers back will be the only practical way to go for remotely serious applications for a while still
It reminds me a bit of genetic algorithms. GA is the 'last resort' when you truly know nothing about how to model your problem.
What is the sweet spot for RL?
I see some sort of reinforcement learning as the most promising technique for overcoming the dramatically named "curse of dimensionality" in the state—the single biggest roadblock to optimizing more complex supply chain models.
In fact, the study of MDPs and their solutions stems from operations research, and I think studying problems in that context give you a powerful way of understanding how reinforcement learning algorithms work. Basic inventory control problems are very intuitive, and there's a natural progression from exact dynamic programming methods (Bellman iteration and policy iteration) to different reinforcement learning algorithms that really helps build an intuition for how RL works.
Could I get in touch with you via PM?
This is most clearly seen when you look at how you would train a supervised learning system to operate an RL agent - you would need to provide the correct action at every timestep. So RL algorithms are mostly interesting when you get periodic reward signals and the reward may depend significantly on actions you did previously, rather than the action you just did. Learning to grip objects is an interesting use case from robotics.
IMO the main reason it's getting more attention is that there is a lot of progress being made, and a lot of that progress is due to progress that is being made in the supervised learning of neural networks.
However, people see some strong parallels between RL and GANs which promise to greatly improve unsupervised and semi-supervised learning. Also there has been work on using RL algorithms (largely REINFORCE) to train non-differentiable parts of neural networks. And then there has been recent work on using RL to decide how to train neural nets over all.
So while most people in industry may never need to touch RL, it will be useful in some systems with time-dependent components and is worth learning from a research perspective.
But there are also a lot of things you can't do. Really any task that requires performing a series of actions to reach some kind of goal. Which covers most of the things we want AI to do. Like controlling a robot, playing a game, talking to a human, proving a theorem, etc.
Of course with regular ML, you could do mimicry, and predict what actions a human would do at every time step. But then you are severly limited by the time and quality of your training data. RL requires no training data and can potentially learn to be much better than humans.
But more generally, when you don't have enough training data ahead of time but do have the benefit of lots of user interactions, and can afford to experiment with live users, then you may be in the sweet spot for RL.
But this is just a characteristic of "online learning" algorithms, no? I thought RL was special method that is online only, but there are other algos that can be made to be online that aren't RL, if my understanding is correct. Then the advantage you cite isn't unique to RL at all.
You can even do online learning with SGD (stochastic gradient descent)
E.g. You ate an apple, you opened the door, you arrived at office, and then had food poisoning.
But I can also appreciate that, from what I'm reading here, that RL brings to center actions/decisions to effect an outcome that might not be as easy to tweak in a supervised setting.
An example here would be a char-RNN. It predicts one character at a time, log probability, and the loss function is the log vs the actual character. Nice and differentiable, so you can take the char-RNN unrolled over 10 timesteps, and at each timestep calculate the gradient to optimize the loss. This also gives you a generative model: sample a character based on the probabilities. Now, take the same char-RNN and redefine the loss as 'whether the user pushed upvote or downvote on the entire 10-character string generated'; you have the unrolled RNN which generated the full string, and you backpropagate... what? What is the gradient for each LSTM parameter, telling it how it should be tweaked to slightly increase/decrease the loss?
RL is also extremely useful in scenarios where rules cannot easily be made explicit. Think of riding a bike or flying a helicopter or gripping objects with a robot arm. Here we can more easily define the reward function - but the agent has to figure out how to do things (to maximise expected reward).
sending the input data to the server, doing the computations there and getting the answers back will be the only practical way to go for remotely serious applications for a while still