Batch Normalization: Accelerating Deep Network Training [pdf]
arxiv.org
arxiv.org
For example, if I make a dataset for the task of distinguishing cocker spaniels from springer spaniels, then when I collect the pictures I can be quite confident about the labels I assign (eg by DNA-testing the actual dogs). But when someone else tries to identify the dogs based only on the pictures I took, they might get some of them wrong.
The 5.1% "human accuracy" number that gets thrown around for ImageNet comes from Andrej Karpathy attempting to manually label part of the ImageNet test set. He wrote a blog post about the experience (http://karpathy.github.io/2014/09/02/what-i-learned-from-com...), which has some interesting details. For example, they did actually find that 1.7% of images were genuinely misannotated: that's because ImageNet annotations come from a consensus of Mechanical Turkers, which is less reliable than DNA testing. :-)
The way the data was collected was not through a 1000-way classification. Instead, we'd query things like "red fox" from a search engine, and then have turkers do the binary task of YES/NO to clean up the results. This is a messy, noisy process but we do our best. In particular, the turkers saying YES/NO to red fox are completely unaware of what other classes are in the dataset - everything they say YES to becomes the "red fox" class in imagenet, in the final discriminative task among the 1000 categories.
Also note that if an image contains a whole bunch of things (e.g. see left-most image here http://karpathy.github.io/assets/ilsvrc2.png), the turker during data collection would be correct in saying that there is in fact, e.g. a "ruler" in that image, but at test time when you're seeing everything in that image, you only have 5 possible predictions so you have to guess about what you think the "true" label is.
Note, the definition of a "human rater accuracy" is a bit vague. For example an "Amazon Turk rater" qualifies as a human rater, yet the accuracy of such rater can be easily exceeded by an accuracy of dedicated grad student rater, or an accuracy of a dedicated H1 employee rater.
edit: That's my guess from previous (2.5 years back) work in computer vision / imagenet and a quick skim over the article.
- 3 human raters, one computer rater
- computer always agrees with 2/3 humans
- humans do not always agree with 2/3 humans
There's some well-funded startup (SkyMind maybe?) who declares on their homepage that deep learning is sooo much better than machine learning (first wtf). Then they explain that neural nets with less than 3 hidden layers is machine learning, with more than 3 layers it's deep learning.(second)
I didn't know if I should laugh or cry.(Not that what they actually seem to be doing is bad, but this newly found hype is just wrong.)
I also mainly advocate neural networks on unstructured data where the results are proven to be a significant improvement over other techniques. Many startups in the space would believe that you can have a simple GUI and you're somehow set to go.
In my upcoming oreilly book Deep Learning: A Practioner's Approach, I go through a practical applications oriented view of neural networks that I think will help change things in the coming years.
I'd also add that I've implemented every neural network architecture and allow people to mix and match. A significant result that many people are familiar with are the image scene description generators done by karpathy et.al[1].
Either way, unlike most our code is open source :). I don't claim things are ideal or perfect, but it's out there for you to play with. I focus on integrations and packaging and providing real value rather than pretending that some algorithm is going to be my edge. Many startups in the space will tell you they have awesome algorithms that are cutting edge when in reality you should be providing a solution for people.
[1]: https://github.com/karpathy/neuraltalk/tree/master/imagernn
I don't see those claims anywhere on your website and I have no idea where the grandparent comment got those statements from. Frankly, they don't make any sense. But I do share his or her sentiment that the phrase "deep learning" is overused, poorly understood and is rapidly becoming a meaningless buzzword akin to "Big Data" or "Web 2.0".
I don't expect you to explain statements that you don't appear to have made, but I would like to hear an expert's view on exactly what deep learning is and how it compares to other machine learning techniques. I understand if you've already addressed this issue in your book and we can read your thoughts on this when it's published. But just a few sentences here might do a lot to clear up some misconceptions for fellow hn readers.
Deep Learning done right applied to unstructured data is a great part of either an ensemble of methods or great for working with hard to engineer features.
I think as for normal machine learning where you are typically doing feature engineering, you need to understand what's going on to make recommendations for actionable insights. Everything has its place.
I'd like to echo Andrew Ng here, the hype in deep learning is overblown. While there are great results, it's not magic.
Neural networks still need feature scaling, among other things to work well. Much of the hype comes from the giant marketing machine that are the PR firms for the research labs who need new data scientists.
While great work is being done in these labs, much of it isn't going to be applicable in the day to day work of a data scientist just trying to do some some basic A/B testing. Hope that makes sense!
Quote:" On a technical level, deep-learning networks are distinguished from the more commonplace single-hidden-layer neural networks by their depth; that is, the number of node layers through which data is passed in a multistep process of pattern recognition. More than three layers, including input and output, is deep learning. Anything less is machine learning. The number of layers affects the complexity of the features that can be identified."
Again, it's great what you are doing in general, I just brought this up as a example of the buzzwords and hype around deep learning.
May I ask what makes a neural network into "not-a-black-box"?
You do mention "proven results". It seems to me that an experiment where one approach does better than another is compatible with one or both approaches being black boxes, ie, there not being more of an explanation than "it works".
But if there's something more going on here, I would love to hear more details.
This could mean debugging neural networks with histograms to make sure the magnitude of your gradient isn't too large, ensuring debugging with renders in the first layer if you're doing vision work to see if the features are learned properly, for neural word embeddings, using TSNE to look at groupings of words to ensure they make sense, or mikolov et. al on word2vec give you an accuracy measure wrt predicted words based on nearest neighbors approaches.
For sound, one thing I was thinking of building was a play back mechanism. With canova as our vectorization lib that converts sound files to arrays, feed that in to a neural network, and then listen to what it reconstructs.
The take away is, while you can't rank features with information gain like random forest, you can at least go in not completely blind.
Remember, one of the key take ways with deep learning is that it works well on unstructured data, aka: things that have brittle feature engineering(manually) to begin with.
Edit: re proven results
Audio: https://gigaom.com/2014/12/18/baidu-claims-deep-learning-bre...
Vision: http://www.33rdsquare.com/2015/02/microsoft-achieves-substan...
Text: http://nlp.stanford.edu/sentiment/
To further clarify what I mean by black box: I don't like typing:
model = new DeepLearning()
What does that mean?
How do I instantiate a combination architecture? Can you create a generic neural net that mixes convolutional and recursive layers? What about recurrent? How do I use different optimization algos?
Not to pick on ml lib, but another example:
val logistic = new LogisticRegressionwithSomeOptimizationAlgorithm()
Why can't it be: val logistic = new Logistic.Builder().withOptimization(..)
The key here is visualization and configuration.
Hope that makes sense!
Perhaps a thing with lots of parameters and some tools/rules-of-thumb for debugging might be called a "gray box" where a "real" statical model with a defined distribution, tests of hypothesis validity and so-forth could be called a "white box".
Secondly, if you'd like to dispute the series of records broken by deep learning across many benchmark datasets, feel free to take that up with Geoff Hinton and Andrew Ng. Skymind is hardly saying anything new when it points to the advances deep learning has made in unsupervised data.
https://news.ycombinator.com/item?id=4779647
There's lots of other results for 'deep networks' in the search too (sorting by date, they appear quite consistently, many of them don't get many votes though).
This same week Microsoft released a similar result, but without such an important new approach for networks.
The important chart really is on page number seven.
The short story is that I tried for 500 images and then got 5.1% error on the test set. Looking at my mistakes, I then tried to differentiate two sources of error: a genuine problem with the dataset (e.g. many correct answers in an image, or incorrect label), and errors I felt could be eliminated by an ensemble of very committed humans who were even better than me at classifying dogs :) And that optimistic error rate is approx 3%.
At test time (once training is finished), they are very efficient though, on orders of milliseconds per image.
Four things pop out at me from the paper:
1) The whitening (per batch) & rescaling (overall) is a neat new idea. But (as referred to in their p5 comment about the bias term being subsumed) this also points to the idea that the (Wu+b) transformation probably has a better-for-learning 'factorization', since their un-scale/re-scale operation on (Wu+b) is mainly taking out such a factor (while also putting in the minibatch accumulation change).
2) The idea that this could replace Dropout as the go-to trick for speeding up learning is pretty worrying (IMHO), since the gains from Dropout seem to be in a 'meta network' direction, rather than a data-dependency direction. Both approaches seem well worth understanding more thoroughly, even though the 'do what works' ML approach might favour leaving Dropout behind.
3) The publication of this paper, so closely behind the new ReLu+ results from Microsoft, seems too coincidental. One has to wonder what other results each of the companies has in their back-pockets so that they can repeatedly steal the crown from each other.
4) For me, the application to MNIST is attention-grabbing enough. While I appreciate that playing with Inception (etc) sexes-up the paper a lot, it raises the hurdle for others who may not have that quantity of hardware to contribute to the (more interesting) project of improving the learning rates of all projects (which is quite possible to do on the MNIST dataset, except that it's pretty much 'solved' with the error cases being pretty questionable for humans too).