Scientists See Promise in Deep-Learning Programs
nytimes.com
nytimes.com
The intuition is let's say I told you to write a complicated computer program. Let's say I told you that you could use routines and subroutines, but you couldn't use subsubroutines, or deeper levels of abstraction. In this restricted case, you could write any computer program, but you would have to use a lot of code-copying. With arbitrary levels of abstraction, you could do code reuse much more elegantly, and your code would be more compact.
Here is a more formal description: If you have a complicated non-linear function, you can describe it similarly to a circuit. If you restrict the depth of the circuit, you can in principle represent any function, but you need a really wide (exponentially wide) circuit. This can lead to overfitting. (Occam's Razor) By comparison, with a deep circuit, you can represent arbitrary functions compactly.
Standard SVMs and random forests can be shown, mathematically, to have a limited number of layers (circuit depth).
It turns out that expressing deep models using neural networks is quite convenient.
I gave an introduction to deep learning in 2009 that describes these intuitions: http://vimeo.com/7977427
Are you sure it's exponential ?
If you look at binary functions (ie. boolean circuits) any such function can be represented by a single layer function whose size is linear in the number of gates of the original function (I think it's 3 or 4 variables per gate) by converting to conjunctive normal form.
Of course it's not obvious that a similar scaling exists for non-binary functions but I'd be a bit surprised if increasing depth led to an exponential gain in representational efficiency.
I am confident, though, based upon my reading of secondary sources written by people that I trust.
From one of Bengio's works (http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf): "More interestingly, there are functions computable with a polynomial-size logic gates circuit of depth k that require exponential size when restricted to depth k − 1 (Hastad, 1986)."
For example, in face recognition, first level - could be pixels. Second level - edges and corners: http://www.cs.nyu.edu/~yann/research/deep/images/ff1.gif Third - parts of the face: http://people.cs.umass.edu/~elm/images/face_feature.jpg
Deep learning, on the other hand, is about using layers of classifiers to progressively recognize higher-order concepts. In computer vision, for example, the first layer of classifiers may be recognizing things like edges, blocks of color, and other simple concepts, while progressive things may be recognizing things like "arm", "desk", or "cat" from the lower-order concepts.
There's a book I read a while ago that was super-interesting and digs in to how one researcher leveraged knowledge about how the human brain works to develop one of these deep learning methods: "On Intelligence" by Jeff Hawkins (http://www.amazon.com/On-Intelligence-Jeff-Hawkins/dp/B000GQ...)
All currently used deep learning algorithms are special cases of neural networks. The reason why this is called "deep" learning is that before 2006, no one knew how to efficiently train neural nets with more than 1 or 2 hidden layers. (Or could, because of computing power.) Thanks to a breakthrough by Dr Hinton, this is now the case.
But all models used are neural nets. It's just that a vast amount new algorithms for training them have been developed in the last years and people came up with new ideas on how to use them.
But it is all neural nets. And that's the whole beauty of it.
edit: I'm not sure I was clear enough-- the term "neural network" is a misnomer that encompasses extremely different models that are largely unrelated except for being vaguely inspired by the brain. A vanilla multilayer-perceptron is essentially a generalization of logistic regression. Restricted Boltzmann Machines are different beasts-- they're a restriction of undirected graphical models made amenable to efficient training. Recurrent neural networks aren't in any way a minor extension of other neural networks-- you need different terminology to talk meaningfully about them and they essentially don't have reliable training algorithms. This latter class can be viewed as Turing-equivalent computation, but they're not at all the same as the models in the original article.
Years later, Swiss researchers (Dan Ciresan et al) found that you can train neural nets with backprop, but you need lots of training time and lots of data. You can only achieve this by making use GPUs, otherwise it would take months.
Check out the publications by Ciresan on MNIST, have a look at Hinton's dropout paper or at the Kaggle competition that used deep nets. Or try it yourself and spend a descent amount of time on hyper parameter tuning. :)
The deep learning / RBM tutorial here is quite good and explains the technique.
It's been awhile since I read the paper, but as I recall it involved training an unsupervised model layer-by-layer (training a layer, freezing the weights, then training another layer on top of it).
As far as I can tell there hasn't been any single revolutionary breakthrough in this field...we just keep getting more computing power, discovering better tricks and heuristics, and trying to build larger and larger networks.
These are all neural nets (with some bells and whistles in some cases like tied weights, pooling units, etc) trained exactly as they were trained before using stochastic gradient descent or LBFGS. We did come up with a lot of tricks for making SGD work though, like momentum terms, clamping of weights during learning, dropout, unsupervised pretraining, etc., but in large part it's just a lot more compute power. These networks just turned out to work very well when you have a LOT of (fairly homogeneous) data and can afford to scale them up computationally. And that's pretty awesome, looks like we have a powerful hammer and there are plenty of nails lying around :)
Hinton has a class on Coursera--I think it would be very confusing for beginners, but it has really great material.
Also, I run the "SF Neural Network Aficionados" meetup in san francisco and will be giving a workshop in January about building your own DBN in python, so feel free to check that out if you're in SF (although space was an issue last time).
1) It's not. It's just a buzz word that people are going to use to separate the current (2006+) research from older research concerning neural networks. This is to draw a clear (and potentially self serving?) distinction between old neural networks that were discredited due to their lack of results vs current research that produces much much better results. So, it's neural networks rebranded. Oh my.
2) "traditional" neural networks and what people are using now are very different--mostly because what people are doing now actually works. Deep learning refers to deep neural networks, which take more traditional neural networks and stack them on top of each other to form a hierarchy of representations that ends up being effective for all kinds of stuff. Not that deep neural networks are a new concept--the newness is more that this is now practical rather than theoretical.
So, deep learning is like 20% bullshit, 80% the real deal. Still lots of work to be done, but I think "deep learning" is a nice buzzword to describe the current state of the art as far as neural networks go. It's all neural networks--but this time it's different, haha.
I also don't recall anyone successfully incorporating the date of the rating into the RBM. Mostly this was useful in other models because on any particular day people would just bias their ratings up or down a bit. But also, as one can imagine, over the course of a year or two their tastes would change. Is it straightforward to include that time dimension into RBMs, and if so, is that a recently discovered technique?
Regarding the DBM - I also tried to use more than one layer, and without success. I tried out 3-layer and 4-layer autoencoders (can be called 1.5-layer and 2-layer DBM), with initialization by stacked RBMs or without it. It did not work well probably because: a) the model was inaccurate, and b) the learning method proposed for DBM was not completely correct. Intuitively, the right DBM-like model with the right learning method should have a chance to improve something on the Netflix task.
I found some improvement though (rather learning time than accuracy) in the standard RBMs. Instead of using CD, I split the weights into two sets, creating a directed RBM version. The "up" weights from the visible nodes to hidden are learned with CD with T=1. The "down" weights are learned to best fit the visible nodes, using the hidden nodes as predictors. The hidden nodes generated by CD T=1 are good enough, and we do not need additional iterations with increased T.
About 12 years ago, I switched from a Bio major to CS. I hoped to major in AI, but after taking 2 upper level classes, one focusing on symbolic AI and the other focusing on Bayesian networks, I was completely turned off.
Our brains are massively parallel redundant systems that share practically nothing in common with modern Von Neumann CPUs. It seemed the only logical approach to AI was to study neurons. Then try to discover the basic functional units that they form in simple biological life forms like insects or worms. Keep reverse engineer brains of higher and higher life forms until we reach human level AI.
Whenever I tried to relate my course material in AI to what was actually going on in a brain, my profs met my questions with disdain and disinterest. I learned more about neurons in my high school AP Bio class than either of my AI classes. In their defense, we've come a long ways, with new tools like MRIs and neural probes.
The answers are all locked up in our heads. It took nature millions of years of natural selection to engineer our brains. If we want to crack this puzzle in our lifetimes, we to copy nature, not reinvent it from scratch. Purely mathematical theories like Bayesian statistics that have no basis in Biological systems might work in specific cases, but are not going to give us strong AI.
Are these new deep learning algorithms for neural networks rooted in biological research? Do we have to necessary tools yet to start reversing engineering the basic functional units of the brain?
More an experiment than anything else, but for anyone who is interested: https://github.com/taliesinb/gorbm. I don't claim there aren't bugs, and there is no documentation.
The consensus I've picked up from AI-specializing friends is that there are a lot of subtle gotchas and tricks (which Hinton and friends know about but don't necessarily advertise) without which RBMs are a non-starter for many problems. Which I suppose is pretty much standard for esoteric machine learning.
Deep learning is not about fixing the residuals of the current chain. Deep learning isn't even about residuals in the first place. It's about (1) finding good representations of your data (aka feature learning), (2) then adding a discriminative model on top and then (3) tuning everything. There is no relation to boosting at all.
Boosting is a clever way of modelling a conditional distribution. The insight behind the success of pre-training is that, for many perceptual tasks, having a good model of the input (rather than the input->output mapping) is key.
I have no delusion that the algorithms that work for training deep networks are anything like what the brain actually does, but I don't care. There are many tasks where deep neural nets are state of the art.
The greedy algorithm bears some resemblance to boosting in its repeated use of the same “weak” learner, but instead of reweighting each data vector to ensure that the next step learns something new, it re- represents it.
I guess that most people however would not think of this interpretation of greedily pretraining deep networks :). (I wonder if mbq had this in mind).
In the same article your point about good models of the input is mentioned, too (only copy&paste a small part of the paragraph):
Unsupervised methods, however, can use very large unlabeled data sets, and each case may be very high-dimensional, thus providing many bits of constraint on a generative model.
The 2006 paper is really an amazing read in my opinion.
ANN-s are overfitted more often than not.
Not C++ or Python, but lua with lots of stuff: torch 7 [[http://www.torch.ch/]]
Here is the link: http://code.google.com/p/cuda-convnet/