Large Scale Distributed Deep Networks (by Jeff Dean et al)
research.google.com
research.google.com
I resubmitted since I'm pretty sure this is of interest to many HN readers. Examples of why it's interesting include:
1. Google's deep learning work is now being used to power Android voice search (http://googleresearch.blogspot.ca/2012/08/speech-recognition... ); and
2. Dean claims that "We are seeing better than human-level performance in some visual tasks," in particular, for the problem of extracting house numbers in photos taken by Google's Street View cars, a job that used to be done by a large team of people (http://www.technologyreview.com/news/429442/google-puts-its-... ).
Can someone explain to me how this is news, given that the handwritten addresses on snail mail envelopes in the US have been OCR'd by neural networks for more than twenty years now?
It's like a game of "Where's Waldo" on freaking crack.
You have literally no idea how complex this stuff is now then do you?
I do have an idea (literally, even) that there are additional problems having to do with extracting the house number images themselves from full-motion video, but that's an image registration problem and not an object recognition problem.
It's important to remember that performance is critical when talking about machine perception. OCR, handwriting recognition, face recognition, etc. can all be done but at what level of accuracy ? At least until very recently machine performance on these tasks has fallen well short of human level abilities.
Last I heard, the percentage of handwritten mail successfully sorted by a machine had reached in the low 80s. The Russian company Parasoft (who also worked on Newton's online handwriting recognition) has been the leader in this field.
So if anyone has any questions on it, I will try to answer.
Do you know of any tutorial that may guide the beginner using DNN, I have no idea how to choose the number of hidden layers and activation functions.
Thanks!
There's no secret sauce for how to choose the number of layers and the activation functions (and anyone that tells you otherwise is lying). It's all application-dependent and the best way to do it is by cross-validation.
Can answer questions about this topic in conjunction with mr. bravura (I'm working with the code referenced in that paper by Dean et al., here at Google, and did grad school at the same place as bravura did his postdoc).
WRT the DNN parameters, is it possible to try all potential possibilities (within reason) and find the best one using only cross-validation, or are there just too many choices and you have to use intuition? (from your comment I don't get if cross-validation is ok to get optimal number of layers etc. or if you have to be "smart")
Thanks for replying!
One of the trends these days is to perform automatic hyper-parameter tuning, especially for cases where a full exploration of hyper-parameters via grid-search would mean a combinatorial explosion of possibilities (and for neural networks you can conceivably explore dozen of hyper-parameters). A friend of mine just got a paper published at NIPS (same conference) on using Bayesian optimization/Gaussian processes for optimizing the hyper-parameters of a model -- http://www.dmi.usherb.ca/~larocheh/publications/gpopt_nips.p.... They get better than state of the art results on a couple of benchmarks, which is neat.
The code is public -- http://www.cs.toronto.edu/~jasper/software.html -- and in python, so you could potentially try it out (runs on EC2, too).
Btw, Geoff Hinton is teaching an introductory neural nets class on Coursera these days, you should check it out, he's a great teacher. Also, you can always come back to Google, we're doing cool stuff with this :)
Thanks!
Topics will be presented in multiple ways (simple and intermediate) so you can have plenty of different views on the same material. The material works as both zero-knowledge intro to the topics as well as quick refreshers if you haven't seen the material in a while (quick -- what's an eigenvector?!).
The launch courses will be 1.) real-world applications of probability and statistics (signal extractions), 2.) linear algebra for computer science, and 3.) wildcard (a random assortment of whatever the heck we think is important or entertaining to know). Future courses are: introduction to neural networks, introduction to computational neuroscience, introduction to deep learning, advanced deep learning, how to take over the world with a few dozen GPUs, avenues by which google will become irrelevant, and robotics for fun and evil.
This is phase zero of a four phase plan. I'll get some pre-launch material together to shove down HN shortly, then it'll launch a few days later. Hopefully you'll hear about the project again.
Sometimes these are multi-layer neural nets, but they don't have to be.
How difficult is it to make a neural net "forget"? If you've already trained over a large set of data, can you un-train it for some of the inputs?
In theory stochastic gradient descent allows you to escape local minima: the noise in the stochastic error surface will likely be enough for the network to escape whatever minimum it is in. Practically, because weights will tend to have a large magnitude and because most of the times you'll be using a saturating non-linearity (such as a sigmoid), the number of steps required to escape that local minimum might be too big.
Presumably, you could use second-order optimization methods to perhaps escape from minima -- because it allows you to make "bigger" steps -- but that comes with its own set of problems (negative curvature being one of them).
I encourage you to actually test these hypotheses: train a simple network on something stupid like MNIST, and make it achieve a reasonable error with many passes through the data. Then change the labels of 10-20-50% of your inputs and continue training (with the same learning rate... or not!) to see how long it takes for the network to get to another minimum.
In a nutshell, you learn a vector of real-valued parameters for each word in your vocabulary. To train a network on sequences of words, you represent said sequence as a concatenation of the vectors of these words, and feed it as an input to the network.
To learn these vectors, you define the problem of "language modeling" as that of discriminating between two sets of sequences: S1 and S2. S1 is the set of sequences that occur in Wikipedia (of which there are many) and S2 is the same as S1, but where you replace a word in each sequence with a randomly chosen word from your vocabulary (which makes it, with very high probability, an invalid sequence of words).
Basically, by learning to discriminative between "good" and "bad" English word sequences, you can learn a language model of sorts. The model is represented by those vectors for each word.
You can then project those vectors into 2D, as bravura did a while ago, and look at what is close to each other: http://www.cs.toronto.edu/~hinton/turian.png
What changed?
As for overfitting the best way to reduce it is to use more training data and I believe the nets discussed in these papers have been trained on some of the largest training sets ever used.
The other problem with deep networks is that they have been considered very difficult to train with backpropagation due to vanishing or exploding gradients. I think the recent major algorithmic developments have mainly involved methods to mitigate these problems.
I struggle with the mathematics used in neural networks. I can understand code but as soon as I start to see calculus my brain freezes over. Does anyone know of a good online course that can give me a crash course in the mathematics required for neural networks?
My CS bachelors covered this, but I was lazy and drank too much beer. Now 20 years later I want to understand it properly.
I have spent a lot of time on Khan Academy to learn the calculus. In my experience you can get by with a surprisingly small amount of calculus, but it happens to be a small amount from a high level.
For example, backpropagation is just repeated application of the chain rule. Did take a while to get a handle on the derivatives, but it's worth it.