Large Scale Deep Learning – Jeff Dean [pdf]
static.googleusercontent.com
static.googleusercontent.com
The best research results from 2014 and 2013 make less use of the unsupervised techniques than initially expected, so I would start by focusing on the below sections, which focus more on supervised learning with deep neural networks:
Sparse Autoencoder: Neural Networks, Backpropagation Algorithm
Building Deep Networks for Classification: Deep Networks: Overview, Fine-tuning Stacked AEs
Working with Large Images: Feature extraction using convolution
You'll need some background in matrix algebra, calculus, and probability to understand this. Having taken a previous machine learning course, although not strictly necessary, is probably extremely helpful--I'd recommend taking any standard course on ML on Coursera or Udacity, or going through any standard textbook.
EDIT: I almost forgot that Michael Nielsen (who wrote the standard textbook on quantum computation) is also writing a free online textbook on Neural Networks and Deep Learning. Chapters 1-4 are currently available and would get you pretty far: http://neuralnetworksanddeeplearning.com/
Michael's book seems to target a more introductory level - a beginner might be better off to start with that, follow with Andrew Ng's ML course, which has a section on neural nets including an assignment implementing backpropagation, then continue with the deep learning book and the {deep learning, UFLDL} tutorials. This should be solid enough to at least read most of the cutting edge work and papers, if that is the aim.
Hugo Larochelle's youtube course https://www.youtube.com/playlist?list=PL6Xpj9I5qXYEcOhn7Tqgh... and Hinton's coursera course https://www.coursera.org/course/neuralnets are also great references.
BTW, for anybody who wants to learn machine learning in general, Kyle's blog also seems to be packed full of clear explanations with working demo code: http://kastnerkyle.github.io/
Very nice!
Pg 26, quote: "Anything humans can do in 0.1 sec, the right big 10-layer network can do too". That is a very bold claim. It encompasses the entire fields of image and voice recognition as well as knowledge encoding. It's slowly becoming clear that this is likely to be true.
Pg 39, 40: Google's ImageNet-winning system in 2011 had 7 layers and an error rate of around 16%. The 2014 system had 24 layers and an error rate of 6.66%. Note that trained humans have an error rate of around 5%[1].
Page 50-57 talk about the miracle that is Word2Vec, and what is possible with that.
Page 60-70 talks about paragraph embedding. I haven't seen this published before.
Page 70-73 extends word/paragraph embedding for translation. I've seen a slide deck showing this works before, but I need to read the new paper cited there.
Page 74+ talks about cross-modal embeddings, especially the caption generation stuff. HN has had a few things on that over the past month or so.
[1] http://karpathy.github.io/2014/09/02/what-i-learned-from-com...
That's an interesting direction from which to view things.
Then AI progress can be measured by increasing that timeframe.
Though I suspect there are some pretty gigantic discontinuities in there. The things humans can do in 3-4 seconds are qualitatively different from what they can do in less than 1 second, for example.
Still, it's a useful perspective to keep in mind.
Paragraph embedding was published this year as "Distributed Representations of Sentences and Documents". http://arxiv.org/abs/1405.4053
It's interesting how the state of the art is outpacing publishing.
From a quick scan that appears quite similar to the approach in papers like "Parsing Natural Scenes and Natural Language with Recursive Neural Networks" (2011)[1]. Edit: I see they cite this paper too.
[1] http://nlp.stanford.edu/pubs/SocherLinNgManning_ICML2011.pdf
Edit: a link about this. https://news.ycombinator.com/item?id=8660624
That being said: I will be benchmarking deeplearning4j's glove with word2vec here soon. Any machine learning algorithm is better when you tune it.
I personally like glove due to having less knobs. The mechanics involving document statistics being part of the gradient update is also interesting.
I've also messed quite a bit with the distributed representations.
I'm not partial to any particular implementation. I'll use what works. That being said, I'm not armchair. I'll be backing this up with my own data as well.
My main point is just because something is controversial shouldn't stop you from trying it. That's what research is: trying new things.
Actually, I'm not sure we can do those things in quite 0.1 sec. All I know about this is from around minute 9 from http://www.radiolab.org/story/267176-never-quite-now/ . One guest on the show even estimates thinking the simplest thought to be on the order of 0.25-0.5 sec.
Actually this was argued by Connectionists in 1980s. It is called Feldman's 100-step rule:
The critical resource that is most obvious is time. Neurons whose basic computational speed is a few milliseconds must be made to account for complex behaviors which are carried out in a few hundred milliseconds (Posner, 1978). This means that entire complex behaviors are carried out in less than a hundred time steps. Current AI and simulation programs require millions of time steps.
Feldman, J. A., & Ballard, D. H. (1982). Connectionist models and their properties. Cognitive Science, 6, p. 206.
> E(hotter) - E(hot) + E(big) ≈ E(bigger)
> E(Rome) - E(Italy) + E(Germany) ≈ E(Berlin)
These things are linearly separable?!
I work with distributed deep nets quite a bit. It's a different animal than training on a GPU.
I am working on benchmarks with my framework deeplearning4j now.
That aside, a few neat references/projects that will be digestible for people.
For those of you already in neural net land, there's a few key takeaways when doing distributed neural nets:
parameter averaging across mini batches
(depending on the algorithm) adagrad
momentum
[1] Project Adam: http://www.wired.com/2014/07/microsoft-adam/ [2]: Associated Paper: https://www.usenix.org/system/files/conference/osdi14/osdi14...
[3]: Hogwild algorithm: http://www.eecs.berkeley.edu/~brecht/papers/hogwildTR.pdf
[4]: A variation of this I use called Iterative Reduce done by my partner Josh Patterson: https://github.com/jpatanooga/KnittingBoar/wiki/Iterative-re...
[5]: Sandblaster LBFGS by Dean and Co. http://research.google.com/archive/large_deep_networks_nips2...
Facebook's AI research lab has contributed to the Torch7 project (which is unsurprising since it is lead by Yann LeCun, and Torch7 was originally developed in his group at NYU).
I wouldn't go as far as to say it's "becoming the industry standard" though. Caffe and Theano are also very popular.
Metacademy is also a very useful resource for anything machine learning: http://www.metacademy.org/
PS: For people who are saying you can "apply" DNNs in a day or learn it by a coursera course in 6 weeks - they are only very superficially right. Yeah, anyone can build ML model for a sample training data using tool in the same sense that anyone can compile sample code and have a working app. The problem is that most models don't work the first time as expected. The challenge lies in debugging the model and fix many of N possibilities to make it work. This is what working in ML is all about. It's like usual programming where it takes years of experience to debug the code and make it work for your purpose. The added twist in ML is that debugging is almost entirely statistical. When your model doesn't work, it doesn't work only in statistical sense. Your problem would be essentially that the model doesn't give expected answer this 12% of the time. For this 12% of the time, it doesn't work not because of some wrong "if" condition or misplaced subroutine call. The debugging is almost always statistical debugging - there are no breakpoints to put or no watch to set or not even exceptions. So it takes pretty solid background in statistics and probability to effectively work in ML. And yes, most likely it would take much more than 2 years.
Watching videos of presentations and reading slides is often much easier than comprehending papers, though ultimately the paper should have much richer detail.
Personal anecdote: Two years ago I just started learning about these things, coming from an undergraduate degree in electrical engineering. Now I am in graduate school for deep learning and AI working to push things forward, one small step at a time. It is totally possible to learn this stuff in a reasonable amount of study, and there are more free resources than ever. Note that I had a full time engineering job until 6 months ago... doing something totally different!
Scaling Deep Learning, Wednesday, December 10th, 2:00PM-3:00 PM at the McGill University M1 amphitheater of the Strathcona building at 3640 University Street.