I found these helpful while researching the history: http://www.andreykurenkov.com/writing/a-brief-history-of-neu... http://www.scholarpedia.org/article/Deep_Learning
What else have you found particularly useful?
I found these helpful while researching the history: http://www.andreykurenkov.com/writing/a-brief-history-of-neu... http://www.scholarpedia.org/article/Deep_Learning
What else have you found particularly useful?
http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf
http://people.idsia.ch/~juergen/deep-learning-conspiracy.htm...
Were I in your shoes, I would NOT have highlighted the Google YouTube experiment as "the" big breakthrough. It was just an interesting worthwhile experiment by one of many groups of talented AI researchers who have made slow progress over decades of hard work. Why single it out?
--
PS. The YouTube experiment did not produce new theory, and from a practical standpoint, it would be unfair to say that it reignited interest in deep learning. Consider that the paper currently has only ~800 citations, according to Google Scholar.[1] For comparison, Krizhevsky et al's paper describing the deep net that won Imagenet (trained on one computer with one GPU) has over 5000 citations.[2] And neither of these experiments deserves to be called "the" big breakthrough.
[1] https://scholar.google.com/citations?view_op=view_citation&h...
[2] https://scholar.google.com/citations?view_op=view_citation&h...
The reigniting of deep learning around 2012 was because of Krizhevsky, Sutskever & Hinton winning the Imagenet challenge (1000 object classes)
Contrary to how much Google tried to sell Andrew Ng's "breakthrough" 2012 experiment with tons of PR, the paper is very weak, and cant be reproduced unless you do a healthy amount of hand-waving. For example, to get an unsupervised cat, you have to initialize your image close to a cat and do gradient descent wrt the input. Or else, you dont get a cat... It is not even considered a good paper, forget being breakthrough. Also, those 16000 CPU cores etc. can be reproduced with a few 2012-class GPUs and much smaller time-span than their training time.
The next slide after the 2012 breakthrough that shows the Javascript neural network -- contrary to what it looks like -- is not TensorFlow either. It has cleverly and conveniently been given TensorFlow branding, so most people just confuse it, but it's just a separate Javascript library akin to convnet.js
Since your page gets a ton of hits, it's at least worth it to publish a comment about these GLARING inaccuracies.
1. They demonstrated a way to detect high level features with unsupervised learning, for the first time. That was the main stated goal of the paper, and they achieved it magnificently.
2. They devised a new type of an autoencoder, which achieved significantly higher accuracy than other methods.
3. They improved the state of the art for the 22k ImageNet classification by 70% (compare to 15% improvement for 1k ImageNet in the Krizhevsky's paper).
4. They managed to scale their model 100 times compared to the largest model of the time - not a trivial task.
You say "it can't be reproduced" and then "can be reproduced", in the same paragraph! :-)
Regarding initializing an input image close to a cat to "get a cat", I think you missed the point of that step - it was just an additional way to verify that the neuron is really detecting a cat. That step was completely optional. The main way to verify their achievement was the histogram showing how the neuron reacts to images with cats in them, and how it reacts to all other images. That histogram is the heart of the paper, not the artificially constructed image of a cat face.
It's not perfect, but I cant give a reasonable answer to the other extreme of an opinion. fwiw, as a researcher I spent quite some time on this paper, but that subjective point doesn't mean anything to you.
The only negative thing I can say about that paper is they have not open-sourced their code.