Understanding LSTM networks
colah.github.io
colah.github.io
Yep! One computes the gradient with backpropagation and trains LSTMs on that.
> It is very odd that SGD on the error function for some training data is conceptually equivalent to teaching all the gates for each hidden feature when to open/close given the next input in a sequence.
Agreed, it's pretty remarkable the things one can learn with gradient descent. I'd like to understand this better.
It's probably a case of my intuitions about optimization being really broken in high-dimensional spaces, but I'd like to improve that. I'd also like to understand how the cost surface we're optimizing on interacts with network architecture decisions. Both, unfortunately, are very hard demands.
This has been the general consensus (based on intuition) until recently. It turns out that it's not as big problem as tought. Almost all local minima have very similar error values. The bigger issue is combinatorially large number of saddle points where gradient becomes zero with very few downward curving directions. Saddle-free Newton method is one way around that.
Open Problem: The landscape of the loss surfaces of multilayer networks, Choromanska, LeCun, Arous http://jmlr.org/proceedings/papers/v40/Choromanska15.pdf
The weights are initialized as noise, the optimization problem isn't convex so there are numerous local optima to get stuck in, and the distance between the relevant input and output can be dozens or hundreds of timesteps in length. The last point is quite important as you can end up with very odd gradients (exploding / vanishing gradients). Troublesome gradients are why traditional RNNs don't work particularly well - they can't actually make the connection between input and output.
Others can likely explain the intricacies better :)
edit: It's kind of like, if not actually equivalent to, programming a Turing machine by its gradient on training data.
Computing the gradient on random mini-batches ends up helping "average things out". But adding more layers, different architectures, different regularization techniques, you get very different results than you would get by taking the simpler mini-batch only approach where my expectation would be more of an "averaging" effect along a smooth surface rather than the kind of learning we see from a neural network.
It's remarkable because it seems you can get away with ignoring certain assumptions about the smoothness of the error surface and bootstrap predictive models from raw (read: cheaper to generate) data rather than features pre-designed to live in a nice space together. A lot of people didn't think you could do that, and the question of why you can remains unanswered.
I hope you keep writing as much as you can. Thanks!
It's really impressive to me how farsighted his work on LSTMs was.
In any case, I'm glad you liked the post.