Deep Learning - The Biggest Data Science Breakthrough of the Decade
oreillynet.com
oreillynet.com
Maybe if you don't have anything on topic to say, just do not comment? You really are not obliged to have an opinion on everything.
(Waiting for the downvotes)
It is also a matter of opinion how widely they are being used in industry. Certainly they are being studied in many companies but they do not a appear to used much in production because of their complexity and high training cost. This is still cutting edge technology.
In my experience most professionally trained mathematicians and statisticians are still pretty skeptical of these claims. Wouldn't you agree?
I would love to see a breakthrough in data cleansing or how about just standardized coding, labeling and formatting. Unfortunately I've used lots of 3rd party data sources and wasted more time on these brainless activities than I want to think about. Consider yourself lucky if you only work on web logs where you control what they look like.
I have degrees in Physics, EE and a PhD in CS. I munched axiomatic set theory and infinite ordinals during the course of my PhD. I dabbled in theoretical machine learning for 3 years. See, I can play the credentials game too.
But does that address the fact that they are not used in the industry? AI is full of charlatans and broken promises. Sadly, by listing "deep learning" alongside Deep Blue and Watson, it seems more charlatany.
[1] http://research.google.com/archive/large_deep_networks_nips2...
2. Model X for intelligence being highly technical or cool in some mathematical way is not scientific validation.
3. Google hiring X is not the same as X's model being successful in the industry. I have long switched over to DuckDuckGo for technical queries.
Anyways, what AI people should first address always is point number 2.
AI has always jumped from one cool thing to the next without answering whether that cool thing has any scientific basis.
Don't bring another AI winter ;)
It is always cool to see excitement over research in AI! (As long as it does not drown out other competitive approaches which might bear fruit in the long run.)
That's exactly what I am complaining about.
2. Model X for intelligence being highly technical or cool in some mathematical way is not scientific validation.
I never said that, it just narrows the amount of people that can comment on it with merit.
3. Google hiring X is not the same as X's model being successful in the industry.
It was just a side-comment.
How is one to scientifically validate against something that can't even be defined?
First of all that is completely wrong even for simple things like image recognition (try building a face recognizer which works under all possible conditions).
But more starkly consider the following question:
Is Geoff Hinton a machine?
Again, you're quoting me on that. Yes, 'not used much in data science' is a valid argument that it's not one of the biggest breakthroughs in data science.
And if you want to discuss the topic (while blanket criticizing people for not knowing what they're talking about) at least get the father of deep network's name right: it's Geoff Hinton, not George.
Sure, I only have a graduate math degree and only follow these latest developments casually and perhaps I just miss exact way this newest artificial neural network stuff is really that different than the older stuff. But the only thing that's being touted is a NYTimes article. As another poster said, if you'd like to add to the conversation, give us some "meat" here.
My small exposure to ML also left me feeling the whole train, test, operate cycle is a pain in the neck.
The main difference between these newer networks (besides much improved performance) is that the algorithms can handle "deeper" networks better (more hidden units). If we're talking about Deep Belief Networks, they're not much like the old ANNs. DBNs are generative probabilistic graphical models using Bayesian inference.
Conceptually, going deeper (LOL) allows the networks to learn higher level concepts. For example, a 1 layer ANN (perceptron) can only learn linear functions, while a deep network is able to internally form a belief of what, say, a cat is.
More technically: Much of the work in ML is deciding what your inputs (features) should be. When classifying text documents, should you use word counts, bag of words, word stemming, character counts, etc. Should the model be linear, polynomial, gaussian, trigometric, etc. Deep learners try to automatically do feature selection and control the degrees of freedom in the model for you.
Also, deep learning is catching on in some industries. It has recently had huge successes in speech recognition, and all major companies developing this technology have started using it (e.g. Siri for one).
https://news.ycombinator.com/item?id=5376319
they use DNNs not DBNs (DBNs only used for pre-training, sometimes). Also if you read Microsoft's paper Table2, 7 hidden layer networks, which clearly qualifies as deep, work just fine with Back-propagation. Just a bigfat-MLP, no preprocessing! but 17.4 Word Error Rate (WER) vs 17.0 WER for DBN pre-training.
>they're not much like the old ANNs. DBNs are generative probabilistic graphical models using Bayesian inference.
MLPs (DNNs) can also be interpreted probabilistically. Just a directed model where inference is attained by marginalization of the hidden binary nodes in a layer-wise manner and by using a naive mean field approximation. All that to say the classic "forward-pass" ;).
Also could you indicate me a source confirming that Siri (Nuance) also switched to DNNs?. I am interested in that.
>Conceptually, going deeper (LOL) allows the networks to learn higher level concepts.
That is the really interesting part!. Now, I have not seen a proof for that. Wondering at individual neurons modeling individual features of e.g. a face or so is also a trend of the 90s and does not count as proof. I said this because it is what I usually hear.
Until now the justifications I saw for multiple layers of perceptrons being suitable for modeling arbitrary high level abstractions are reduced to
1) MLPs are universal approximators. This in my opinion is a superficial argument. GMMs also allow modeling "any" distribution and Taylor series any linear function, but in reality there are physical limitations to this argument. Maybe is true if you had a billion layer net, but will you get there?. If you had that computing power maybe a more realistic modeling of the brain might work better
2) They resemble how brain architecture works and similar arguments. Which I am fairly sure is not true. There are more human-brain based approaches to AI like e.g. cortical learning algorithms and those just seem to model that stuff to a certain extent.
Well, that ones about me. Yes, I have a plenty of experience in machine learning, including undergraduate research in neural networks, a graduate degree in machine learning, and more than five years of industry experience (including several years building some of the most utilized neural network models in industry). I have read many of the deep network papers in detail, and have played around with them on actual data.
And yes, I think your comment deserves to be downvoted; unlike those of us with insight into the issue you added nothing to the discussion other than derision. It bothers me that comments like yours end up at the top of so many threads like this.
edit: and I'd like to point out as someone in the industry I have a good reason to temper expectations. Undeserved hype leads to bubbles, and bubbles create collateral damage when they pop. The AI industry has dealt with this at least twice already, and I don't want to see it happen again. The results so far are extremely exciting, but deep networks still need to prove they deserve the hype.
By the way, I am not asking anyone what their degree is, I am just asking to bring arguments and experiences or stay neutral.
There isn't great library support for deep networks. This is a big deal, I don't want to spend tons of time building my own library or working with buggy/poorly supported/infant libraries. In production systems we prefer extremely well-established libraries that work in our language/environment of choice. Also deep belief networks are a couple orders of magnitude slower than linear models (probably the most commonly used type of model in industry). They require more parameter selection. They're not even useful for a lot of tasks - if I'm already spending tons of time building useful features (often a requirement in industry for non-technical reasons, like reporting or legal constraints) deep networks aren't going to be very useful. Much of their utility is taking raw, unstructured data and creating useful features for a supervised model. You can't easily interpret them as models unless you are working with visual data, they are a black box.
They say absolutely nothing of relevance other than how awesome it's supposed to be, and oh by the way this is Cloudera and it's great, and I happen to work in Kaggle and it's magnificient. After preying your personal data to let you listen to the infomercials.
I do not know well other examples beyond case of Automatic Speech Recognition, but since this case caused a lot of noise, I bet it is responsible for a reasonable chunk of the Deep learning "buzz". Here is my take about this.
If you look at papers from Microsoft like Seide et al 2011 and similar papers the reported improvement against state of the art (up to 30%) is really impressive and seems solid. Now, the technique is more or less using a very big multi-layer perceptron (MLP), a technique already established two decades ago (or more). There is some fancy stuff like the deep belief network based initialization, but it does not make big differences. The core of the recipe itself is not very new. What has changed is the scale of data we have available and the size of the models that we can handle.
With this I am not implying that this is not a very interesting discovery. But it is important to bear in mind that the change in the amount of data could also make other 20 year old techniques interesting again. On the other hand, neural networks had a bad name in the last years for understandable reasons. They are a blackbox, or at least less transparent than the statistical methods. This makes them prone to cause the "black box delusion" effect. You hear a new algorithm is in town, it has fancy stuff like remotely resembling human thinking architectures or cool math but you can not completely grasp it guts, then "voila!" suddenly you are overestimating its relevance and scope of applicably. MLPs were hailed as "the" tool for machine learning already once, I think for these same reasons. For me the right position here is a prudent skepticism.
On the other hand, this should also push people to try new/old radical stuff since the rules of the game seem to be changing, it is not a moment to be conservative in ML research :).
from the NYT article [1]: "The achievement was particularly impressive because the team decided to enter the contest at the last minute and designed its software with no specific knowledge about how the molecules bind to their targets. The students were also working with a relatively small set of data; neural nets typically perform well only with very large ones."
NNs in general have enjoyed lots of successful practical (commercial) applications in pattern recognition though they were sort of replaced in the "state-of-the-art" by SVMs in many cases until RBMs and DBNs came along. I agree with your caution for skepticism though, only time will tell how good DBNs are.
I think the black box criticism is BS for the most part. In some cases (google's search being a famous example) it might be great to have a human readable and tweakable solution (assuming you have the resources) but for something like recognising handwritten digits from images, not so much.
[1] http://www.nytimes.com/2012/11/24/science/scientists-see-adv...
Agree, but with black-box I meant not something that is opaque to my grand-mother but partially opaque to engineers that implement MLP machine learning applications and the tech-lead that takes the decisions. The thing is that even research people (or maybe specially them) tend to positively bias things they do not completely understand (so I think, maybe its just me ;)). That is what I meant with black-box delusion. As you say only time will tell.
Regarding DBNs, again, the case of ASR uses DNNs which is to say big-fat MLPs. The model is handled as a DBN only for pre-training, and layer-wise pre-training does a similar job anyway.
Any sufficiently advanced technology is indistinguishable from magic, and who knows what wonders magic might accomplish? But once you understand the "trick", it's obvious that it can't do much more than what it's doing. Oh, well. The magic is gone.
There are nice and more balanced overviews here:
http://ufldl.stanford.edu/wiki/index.php/Deep_Networks:_Over...
http://research.microsoft.com/apps/pubs/default.aspx?id=1531...
As I said I can only speak with more or less certainty regarding ASR. I am fairly sure that the success in ASR (with Google and MS embracing DNNs for ASR) contribute significantly to the mainstream impact of deep learning.
http://research.microsoft.com/pubs/157341/FeatureEngineering...
The only reference to differences I found is about differences between a DNN and a MaxEnt models, which is again not an argument for differences between DNNs and MLPs.
Could you point me to a concrete paragraph?, I would be happy to be mistaken in this regard.
I describe some of the key differences between DNNs and MLPs in the webinar. Also, the webinar explains how recent advances go far beyond just applications to speech recognition - in particular I focus on a case study in chemoinformatics.
Agree, as explained in Hinton et al 2006.
http://www.cs.toronto.edu/~hinton/absps/ncfast.pdf
But this is just for pre-training, as I said. If you look at Seides paper, they pre-train treating the MLP as a DBN and then they train it as a classic MLP with BP. Also using layer-wise BP pre-training does bring performance close to DBN pre-training, with no use of DBNs paradigms at all.
>Their structure and training is very different to traditional MLPs
I insist if we are talking of the same DNNs explained in Microsofts paper, this is not true. If we were to be talking about different DNNs please elaborate I would love to hear about that (seriously, no irony here).
http://en.wikipedia.org/wiki/Autoencoder
I am not very familiar with speech recognition, but I think what they talk about here:
Instead of factorizing the networks, e.g., into a monophone and a context-dependent part [5], or decomposing them hierarchically [6], CD-DNN-HMMs directly model tied context-dependent states (senones). This had long been considered ineffective, until [1] showed that it works and yields large error reductions for deep networks.
might be related to this fact. 20 years ago it wasn't known why would you pick a deep network instead of a shallow one, there was even this famous theorem of Kolmogorow that a lot of people in ML misunderstood, that a network with just one hidden layer can in theory learn any function with arbitrary precision.
The problem with NNs is the difficulty of training them. Back propagation with random initial weights is simple, but it can easily converge on suboptimal local maximum if the learning rate is too aggressive. On the other hand, a slow learning rate requires an exponential increase in training time and data. Back propagation as a method was never really broken, it simply wasn't efficient enough to be effective in most situations. Deep belief techniques seem to remedy these inefficiencies in a significant way, while remaining a generalized solution.
Essentially deep belief networks seem to optimize NNs to the point where new problems are now approachable, and greatly improve the performance of current NN solvable problems. The complaint that "the core of the recipe itself is not very new", seems irrelevant in light of the results.
Sure there is. For example, they will never solve the halting problem. They will also (probably) never solve NP-complete problems for very large instances.
I'm sure that a decision tree can also be viewed as a [universal approximator](http://en.wikipedia.org/wiki/Universal_approximation_theorem) if you let tree height go to infinity (just as you need to let layer size grow unbounded with a NN). In practice, this power is at best irrelevant and often actually a liability (you have to control model complexity to prevent overfitting/memorization).
And, importantly, being able to theoretically encode any function within your model is not the same as having a robust learning algorithm that will actually infer those particular weights from a sample of input/output data.
It must be the association with the human brain that just makes neural networks more exciting than other techniques. But dispite the appeal of imitating nature has this usually been the easiest way to make progress in the past? Seems like it would be harder to achieve both goals at the same time.
So far the results are looking pretty good but it is probably best to keep the hype at a reasonable level unless it is crucial of your business model. ;)
And before 1980s style neural networks there were 1950s perceptrons. That was a much bigger mess, it took more than ten years for someone to point out how 'dumb' perceptrons were (they couldn't even model an XOR), which led to a collapse in AI funding that lasted more than 25 years.
You would think that since it already happened with neural networks before it would be less likely to happen again. However it may be that the same factors that lead to the last cycle are still in operation and it is actually more like to happen again. Something like the reasons for the seemingly endless series of real estate bubbles.
While I don't have the math credentials to match Hinton I think as more 'normal' folks like me get into the game there will also be some interesting things going on. We are trying some interesting things that seem very promising, and I'm sure there are lots of other folks beginning to play with these things that will have some interesting ideas and approaches as well.
So I personally think this is super exciting, and while it might not be applicable for every problem Deep Learning will definitely have a big impact.
The "Biggest Data Science Breakthrough of the Decade" in the title is a rather bold claim, I know... But I think it might be justified. If there are are bigger breakthroughs, I'd be interested in people's thoughts about what they might be.
But if you mean the past 10 years I would have to say that the "distributed storage and processing" revolution (Hadoop and others) has had a much bigger impact on data science than all of neural networks including deep networks.
Why the need to hype what is already a well publicized development? I'm starting to cringe whenever I hear "data science" or "big data" and I love this stuff.
"Scaling deep learning to 10,000 cores and beyond":
Presentation Univ. of Washington (March 14, 2013)
https://www.cs.washington.edu/htbin-post/mvis/mvis?ID=1338
You can see one of Quoc's previous talks online:"Tera-scale deep learning: - Quoc V. Le from ML Lunch @ CMU http://vimeo.com/52332329
You may remember Jeff Dean's (http://research.google.com/pubs/jeff.html) post on this: https://plus.google.com/118227548810368513262/posts/PozFb134...
The corresponding research at Google...
"Building high-level features using large scale unsupervised learning"
http://research.google.com/pubs/pub38115.html
http://research.google.com/archive/unsupervised_icml2012.htm...
Previous HN discussion: https://news.ycombinator.com/item?id=4145558
--
How Many Computers to Identify a Cat? 16,000 http://www.nytimes.com/2012/06/26/technology/in-a-big-networ...;
Also, data science involves a lot more than building predictive models. In my experience >95% of effort goes into something other than building a model. In kaggle contests you usually concentrate on that <5%, which IMO is the fun part but it's not the reality of industry. There are many big breakthroughs in data science that don't involve model building.
edit: I haven't listened to the podcast yet (at work), my comment is more about the title.
* Problem definition
* Infrastructure
* Data transformation
* Exploratory analysis (arguably part of model work)
* Results presentation
Then again, this is an ongoing disagreement I have with the Kaggle folks over what constitutes "data science," where I'm pretty confident that "applied machine learning" is a better explanation of what their contests are about.
https://news.ycombinator.com/item?id=4655927
BTW, I'm a big fan of the data analysis that came out of okcupid, is that all your work?
I'd say data transformation is a part of feature engineering (commonly the bulk of the effort in a ML application). And exploratory analysis is part of model work. W/o those 2 one would be building a model out of dreams and wishes.
Data Science is probably a poorly chosen description. I'd say common use includes infrastructure work which for most of us consists in engineering work.
http://research.microsoft.com/en-us/news/features/speechreco... http://research.microsoft.com/en-us/projects/mavis/
That might just be a single application but if it extends into other domains it might end up being very valuable indeed.
It will eventually become just another tool of course just like anything else but if it brings 10-20% improvements in even a few other long-stagnant areas I would agree with saying that it is a big deal.
I don't suppose someone has an alternative version somewhere?
That said, AI, like every other sciences, experiences trends and bubbles. If you give a decent look to usages and problem solving with machine learning, deep learning techniques are not exactly the final answer. Typically, they're slow to train, to my knowledge there is no good 'online' algorithm yet to train them (i.e. for autoencoders, recursive autoencoders, Boltzman machines). Many applications, and a trend toward 'lifelong learning'[1], require fast incremental learning that yields results in near real-time, or at least in minutes rather than days.
I've compared a couple of unsupervised machine learning algorithms with recursive autoencoders: the latter can learn deeper, very often, but at a computational costs (days vs seconds). Deep learning computation will improve, for sure though.
It took me a while to understand that the slides were in a popup was blocked. After reloading, the slides don't match the audio.
I haven't received any marketing stuff from Cloudera or O'Reilly however. Honestly, I doubt those companies would do anything questionable with registrations.