So if anyone has any questions on it, I will try to answer.
So if anyone has any questions on it, I will try to answer.
Do you know of any tutorial that may guide the beginner using DNN, I have no idea how to choose the number of hidden layers and activation functions.
Thanks!
There's no secret sauce for how to choose the number of layers and the activation functions (and anyone that tells you otherwise is lying). It's all application-dependent and the best way to do it is by cross-validation.
Can answer questions about this topic in conjunction with mr. bravura (I'm working with the code referenced in that paper by Dean et al., here at Google, and did grad school at the same place as bravura did his postdoc).
WRT the DNN parameters, is it possible to try all potential possibilities (within reason) and find the best one using only cross-validation, or are there just too many choices and you have to use intuition? (from your comment I don't get if cross-validation is ok to get optimal number of layers etc. or if you have to be "smart")
Thanks for replying!
One of the trends these days is to perform automatic hyper-parameter tuning, especially for cases where a full exploration of hyper-parameters via grid-search would mean a combinatorial explosion of possibilities (and for neural networks you can conceivably explore dozen of hyper-parameters). A friend of mine just got a paper published at NIPS (same conference) on using Bayesian optimization/Gaussian processes for optimizing the hyper-parameters of a model -- http://www.dmi.usherb.ca/~larocheh/publications/gpopt_nips.p.... They get better than state of the art results on a couple of benchmarks, which is neat.
The code is public -- http://www.cs.toronto.edu/~jasper/software.html -- and in python, so you could potentially try it out (runs on EC2, too).
Btw, Geoff Hinton is teaching an introductory neural nets class on Coursera these days, you should check it out, he's a great teacher. Also, you can always come back to Google, we're doing cool stuff with this :)
Thanks!
Topics will be presented in multiple ways (simple and intermediate) so you can have plenty of different views on the same material. The material works as both zero-knowledge intro to the topics as well as quick refreshers if you haven't seen the material in a while (quick -- what's an eigenvector?!).
The launch courses will be 1.) real-world applications of probability and statistics (signal extractions), 2.) linear algebra for computer science, and 3.) wildcard (a random assortment of whatever the heck we think is important or entertaining to know). Future courses are: introduction to neural networks, introduction to computational neuroscience, introduction to deep learning, advanced deep learning, how to take over the world with a few dozen GPUs, avenues by which google will become irrelevant, and robotics for fun and evil.
This is phase zero of a four phase plan. I'll get some pre-launch material together to shove down HN shortly, then it'll launch a few days later. Hopefully you'll hear about the project again.
In a nutshell, you learn a vector of real-valued parameters for each word in your vocabulary. To train a network on sequences of words, you represent said sequence as a concatenation of the vectors of these words, and feed it as an input to the network.
To learn these vectors, you define the problem of "language modeling" as that of discriminating between two sets of sequences: S1 and S2. S1 is the set of sequences that occur in Wikipedia (of which there are many) and S2 is the same as S1, but where you replace a word in each sequence with a randomly chosen word from your vocabulary (which makes it, with very high probability, an invalid sequence of words).
Basically, by learning to discriminative between "good" and "bad" English word sequences, you can learn a language model of sorts. The model is represented by those vectors for each word.
You can then project those vectors into 2D, as bravura did a while ago, and look at what is close to each other: http://www.cs.toronto.edu/~hinton/turian.png
What changed?
As for overfitting the best way to reduce it is to use more training data and I believe the nets discussed in these papers have been trained on some of the largest training sets ever used.
The other problem with deep networks is that they have been considered very difficult to train with backpropagation due to vanishing or exploding gradients. I think the recent major algorithmic developments have mainly involved methods to mitigate these problems.
Sometimes these are multi-layer neural nets, but they don't have to be.
How difficult is it to make a neural net "forget"? If you've already trained over a large set of data, can you un-train it for some of the inputs?
In theory stochastic gradient descent allows you to escape local minima: the noise in the stochastic error surface will likely be enough for the network to escape whatever minimum it is in. Practically, because weights will tend to have a large magnitude and because most of the times you'll be using a saturating non-linearity (such as a sigmoid), the number of steps required to escape that local minimum might be too big.
Presumably, you could use second-order optimization methods to perhaps escape from minima -- because it allows you to make "bigger" steps -- but that comes with its own set of problems (negative curvature being one of them).
I encourage you to actually test these hypotheses: train a simple network on something stupid like MNIST, and make it achieve a reasonable error with many passes through the data. Then change the labels of 10-20-50% of your inputs and continue training (with the same learning rate... or not!) to see how long it takes for the network to get to another minimum.