I'm no AI/ML expert, but I can't believe this is true... Is it?
I'm no AI/ML expert, but I can't believe this is true... Is it?
Just like how we can know somewhat how neurons work from a modeling perspective, but when you bundle millions of them together, what exactly each one is doing is not quite clear.
Why it is that ASGD and backprop converges on a non-convex optimization problem, and what kinds of model topologies make it do better / worse? That's all basically art right now.
But no one understands what an actual trained neural net is doing. You can look at the weights, and you can watch the inputs and outputs, but it is very difficult to understand why it does what it does.
It just fit a model to data, but there is no explanation why that model is best. The weights are not interprettable by humans.
There have been some attempts at making models which humans can interpret. One program named Eureqa fit the simplest possible mathematical expression possible to a set of samples. A biologist tried it on his data and found that it created an expression which actually fit the data really well. But he couldn't publish it because he couldn't explain why it worked. It just did. But there was no understanding, no explanation.
There aren't many details. I believe the scientist who said it is Dr. Gurol Suel, from that radiolab page.
I would phrase that differently: we know when gradient descent and back-propagation work, not why they work for so many real world problems.
For the when, there are zillions of published mathematical results stating "if a problem has property X, this and this method will find a solution with property Y in time T", in zillions of variations (Y can be the true optimum, a value within x% of the real optimum, the true optimum 'most of the time', etc. T can be 'eventually', 'after O(n^3) iterations', 'always', etc)
However, for most real-world problems we do not know whether they have property X, or even how to go about arguing that it is likely they have property X, other than the somewhat circular "algorithm A seems to work well for it, and we know it works for problems of type X"
Perhaps what you're trying to say is we don't know why finding local minima of these problems is good at solving the problem?
So whenever you read about some 50-layer net trained with this architecture with this padding/stride/normalization, and you wonder how they came up with that, the answer is: some grad student sat there, thought about his past experiences and the papers he's read of architectures that have worked well, and then spent months trying a bunch of things.
Yes it's true that hyperparameters seem arbitrary, but that's just a consequence of the no-free-lunch theorem. Some models will always fit some problems better than others. There is no such thing as a perfect model. NNs turn out to be a really good prior for real world problems. I wrote about why that might be here: http://houshalter.tumblr.com/post/120134087595/approximating...
But in principle, that doesn't apply to things like layer numbers. In theory the best neural net has infinite layers of infinite size and infinite convolution, with a stride of 1. Because you can fit every other model into that, and as long as the parameters are properly regularized it can also avoid overfitting too. In the real world of course, we are limited to merely 50 layers, and need to cut corners with convolution sizes and such.
Likewise for training. The best algorithm for training is bayesian inference on the parameters. Since that is usually very impractical, we use approximations like maximum-likelihood or dropout.
It goes without saying that the reason we have 50 (vs. deeper) layer nets and strides, among other things, is not solely to reduce computational cost.
In order to do that you do need to use multiple layers. And the same is true for digital circuit, which NNs basically are. I'm sure there is mathematical theory and literature on the representation power of digital circuits.
There is a limit to what you can do with only one layer of circuits, and you can do more functions more efficiently with more layers. That is, taking the results of some operations, then doing more operations on those results. Composing functions. As opposed to just memorizing a lookup table, which is inefficient.
That's why multiple layers work better. It isn't some strange mystery.
>an infinite number of layers and convolutions makes no sense - I think you meant to use the word arbitrary
A better way to word it would be "as it approaches infinity" or "in the limit" or something. That is, the accuracy of the neural net should increase and only increase as you increase the number of layers and units (provided you have proper regularization/priors.) Since bigger models can emulate smaller models, but not vice versa.
The forward pass of a net is not theoretically interesting. It's the training of the net that has no theory. The training has nothing to do with digital circuits.
It goes without saying that you've handwaved some (perfectly fine) ideas about composing functions and such. And then claim "it isn't some strange mystery." That's my point. You've argued some ideas from intuition. There is little theoretical rigor around this, however.
Not with proper priors/regularization.
>You've argued some ideas from intuition. There is little theoretical rigor around this, however.
There's this paper which goes more into theoretical depth on the idea: http://arxiv.org/abs/1412.0233
While we understand the math of backpropegation, we don't really understand why it lets neural networks converge as well as they do. If you asked a mathematician they'd point out that yes, backpropegation can work, but there are no guarantees even around the probability that it will work. And yet, it does work, and often enough and in enough different cases for it to be useful.
Yes, interpretability can be an issue, but that's not what is being discussed here.
There have been some attempts at making models which humans can interpret. One program named Eureqa fit the simplest possible mathematical expression possible to a set of samples. A biologist tried it on his data and found that it created an expression which actually fit the data really well. But he couldn't publish it because he couldn't explain why it worked. It just did. But there was no understanding, no explanation.
This is a pretty well understood issue in the ML community. Different tools have different strengths: Regression (sometimes), Tree based classifiers and Bayesian nets are nice for explaining your data, but others methods (SVMs, Neural Networks etc) can be better for prediction in some circumstances.
I'm not certain, but for one, stochastic gradient descent is usually used, which helps break out of local optimas. And second, they don't seem to be as much of a problem as you add more hidden units. I was under the impression there was a mathematical explanation for both of those, though I haven't done much research.
Off the top of my head, I think as you add infinite hidden units, a subset of them will fit the function exactly just by chance, and gradient descent will only increase their parameters. In fact as long as they are in some large possibility space of the correct solution, GD can move the parameters downhill to the optimal solution. Not that we even want the optimal solution, just a good enough one.
>The number of local minima outside that band diminishes exponentially with the size of the network. We empirically verify that the mathematical model exhibits similar behavior as the computer simulations, despite the presence of high dependencies in real networks. We conjecture that both simulated annealing and SGD converge to the band of low critical points, and that all critical points found there are local minima of high quality measured by the test error. This emphasizes a major difference between large- and small-size networks where for the latter poor quality local minima have non-zero probability of being recovered.
> Finally, we prove that recovering the global minimum becomes harder as the network size increases
In [1] he talks a lot about how there is no theoretical basis to think that a deep neural network should converge, and prior to around 2006 the accepted wisdom was that networks deep enough to outperform other methods of machine learning were useless because they couldn't be trained. Then they discovered that using larger random values for initialization means that they do converge, but that there is no theoretical basis to explain this at all.
[1] http://www.thetalkingmachines.com/blog/2015/1/15/machine-lea...
[2] https://scholar.google.com.au/citations?user=x04W_mMAAAAJ&hl...
As for why nets converge at all, there's this paper whichI believe tries to establish some theory why bigger nets don't have such a big problem with local optima: http://arxiv.org/abs/1412.0233
Just the same as you could run a high-degree polynomial regression on some data and "not really know" why it chose the coefficients it did or why it yields a particular output for another input, neural networks typically involve a lot of mysterious coefficients, and it's hard to qualitatively explain why the neural network gives a particular output for another input.
But in both cases, the underlying principles and mechanism of action are very well understood.
Indeed, many recent systems rely on CNN-extracted features being passed to a SVN classifier to classify images that the CNN wasn't trained on (because the SVN is easier to train will less data, and using the CNN frees one from doing feature engineering).
1. They don't usually try to explain the specific activations in the hidden layers of the network. This is hard and depends on the specific trained net.
2. They can't guarantee a net's performance before experimenting with it.