Don't use deep learning when your data isn't that big
simplystatistics.org
simplystatistics.org
The author explores sample sizes up to 85, and then suggests this is the relevant range except at Google, Amazon, Facebook, etc.
But the VAST majority of people considering deep learning have sample sizes between those extremes. Results on small samples are interesting, but it's disingenuous to market this as typical of the world outside Google.
This is a very bad argument for the given clickbaity headline. A methodology that works well for one dataset with few observations will not necessarily work well for another dataset.
You can do almost whatever you want with small datasets, it's just harder than with big data (and is necessary if obtaining data is expensive, e.g. medical trials). Specifically, you'll want to do bootstrapping to simulate additional data and reduce the uncertainty due to a low amount of data.
The "almost" is that you can't have hundreds of features if you have a small dataset (Curse of Dimensionality: https://en.wikipedia.org/wiki/Curse_of_dimensionality)
Most of the wins under the "deep learning" umbrella come from extracting meaning from homogenous features like "the pixel at x-2,y+1 has red=123" or "the word at n+1 is 'king'". That's why we see latent variable embeddings like word2vec come from the DL world even though they're not deep.
When you want to include highly informative features in a deep network, it's often better to feed them into a separate logistic model, as shown in the Tensorflow wide-deep tutorials.
State of the art performance on MNIST is held by a 6 layer convnet (4 layers convolutional, 2 layers FC). MNIST is just 28 x 28 grayscale images, so 768 dimensions. There are many more datasets on the same order of dimensionality. CIFAR 10/100 (32 x 32 pixel images) is also dominated by DL convnets, AFAIK.
Data augmentation is also a thing.
Even taking data size out of the picture, functionally it is not there yet for most tasks. Maybe it will be in the future, but the big problem with it is that with n neurons, you have n^n possible topologies, and finding the right neural topology is a major optimization problem that we're only barely learning basic human heuristics for.
I'm willing to bet the deep learning thing is just one more Neat fad that will eventually cause disillusionment at its lack of results, reverting us back to the Scruffy view that intelligence is far too complex to be described holistically by small sets of simple algorithms. The great thing about the Scruffy philosophy is that it isn't derogatory...deep learning will always have a place as a tool in its tool set. It merely doesn't hold unreasonable expectations.
Deep learning is already a Scruffy fad. It basically says, "Hey, let's use a really huge hypothesis space of circuits that often includes a heavy prior towards convolutions." Gradient descent is a Neat principle, but the whole point of things like improved training methods, new objective functions, and convolutions was to deal with the exploding-gradient problem.
Deep learning didn't come up with its own Neat principle, it invented Scruffy methods to apply a Neat principle to a really fucking huge hypothesis space, so long as you've got a pretty big dataset.
I suspect the reason why deep learning has done so poorly in my domain is that the underlying data is a result of things that are very poorly abstracted as a "function". We have lots of discrete events, stateful buffering, hard non-linearities, discontinuities, numerical bounds, etc. It's more like learning business rules and physical process design than learning a mathematical function. This is part of the reason I don't see deep learning being a holistic solution for self driving cars...once you get past sensory perception and simple 2d path planning, driving is more of a rule based process than anything.
That being said, ML tends to be a pretty niche technique for us anyway. If a process and its components are well known and understood, we tend towards solutions that come from Operations Research over Machine Learning. It is only when things are poorly understood that we use ML (example: predicting product demand fluctuations based on media coverage or predicting truck arrival times given severe weather patterns and traffic backups). PGMs do really well here, but are far more difficult to understand, formulate, and train...for most tasks Random Forests are almost always Good Enough(TM).
Understandable models with clear intervention points are what most businesses seem to need once you get to digging around in their operations, customer and sales data.
1. http://www.openias.org/hybrid-generative-discriminative
2. https://pdfs.semanticscholar.org/b6b9/39ffc9920cd8521299a6fe...
That said you can do a lot with a relatively little set. This 2012 paper puts the range between 80-570 samples [1] again depending on model and required outcomes. Leslie Smith at NRL has been working on this problem as well and showing some great progress on really small sample sets as well.
Major takeaway here is that there is such a thing as too big, and too small of data sets for classification accuracy, but those definitions are rapidly changing.
^Your mileage may vary depending on model, fine tuning, transfer learning etc...
I don't understand how this could be true. Shouldn't the sweet spot be a function of the dimensionality of the data?
If DL would need more training data for higher dimensional inputs, then DL would lose against a simple pattern matching (correlation) algorithm at some point.
Then imagine classifying the color of a single pixel as "light" or "dark". There's three dimensions-- red, green, and blue. You would also need much less training data here than if you were trying to train a network to recognize a car, right?
I think this is what zeroxfe is referring to
The major advantage of deep learning is not that it works better on more data. It's that it automatically learns features that would otherwise take expert humans a lot of time and energy to figure out and hardcode into the system.
This meant that you could train on large unlabelled data and then small amounts of labelled data.
1. Straw man tweet by some non-practitioner which is used to set up the straw-man argument.
2. The whole Digits example is ridiculous, statisticians "love" toy problems to prove theorems & make "arguments" etc. ML is empirical and not just the performance but the entire pipeline from data to application matters.
Let me illustrate: If your aim is to predict 1 vs 0 from images of digits. As an ML researcher I would write a program to synthesize images in all different combinations of fonts/font-color/background color/ location available. The data would easily be more than ~100,000 images. At this point one cannot use LASSO on top 10 pixels (due to jittering), and a Deep Models would be necessary. But in reality my model will outperforms because the thinking process as an ML researcher was not to make an "argument" but to "solve" the problem of detecting 1 vs 0.
3. But the biggest flaw is the following argument """The sample size matters. If you are Google, Amazon, or Facebook and have near infinite data it makes sense to deep learn."""
This is an another issue with Biostatisticians (The author of this post is Bio-Stats professor), is that they are fundamentally unable to recognize importance of programming and ability to collect data. Even if you are not Google, Amazon, Facebook you can easily collect data, even labeled data in scale of terabytes can be collected in within days or a week. Every single PhD student I know is not limited by size of the data but rather computational power and storage available to them. I personally have several terabytes of video and data from YFCC 100M that I would love to process and build models on but I am only limited by the computational power & AWS costs. If you want a concrete example, see the Google PlaNet paper [1] I today have enough data (~5 Tb) to replicate it and build open source geolocation model, the only hurdles are storage and computation costs.
How much of this is students doing research where they already have access to big data, which makes sense if your goal is to do deep learning research, vs being given a problem a business wants to solve? Can you make the same statement for the average problem at your average small-medium sized business? Can you really get big data that is relevant to the local, non-chain coffee shop down the street?
If you can it seems like an amazing business opportunity - to bring Google level insights to businesses that don't directly have Google-level data.
There's a classical quote from Tukey "The combination of some data and an aching desire for an answer does not ensure that a reasonable answer can be extracted from a given body of data." - yes, it's quite likely that an average small-medium sized business has no problems where the possible benefit of ML-driven insights won't match the costs required to analyze whatever data they have.
However, if a small-medium business has some problem with a large enough likely payback to justify making some ML system, it is quite likely that deep learning may be applicable on their data.
A big issue is transfer learning - in many domains while you may have a small amount of data, you'd want a system that has learned to generalize on a huge quantity of similar external data, and just tuned on your data. For example, if a cookie bakery needs analysis of cookie pictures or reviews of cookies, and has limited data samples, it would be reasonable to include e.g. ImageNet data or Amazon review corpus. You'd "teach" the system how pictures/internet reviews/English language/whatever else works on the biggest data available, and just retrain/adapt it to your particular problem afterwards.