Critique of Paper by “Deep Learning Conspiracy”
people.idsia.ch
people.idsia.ch
And to add onto this, another academic is getting his jimmies rustled because he didn't get the money his former PhD students did??
Heavens to Betsy!
---
Everyone self promotes and does things in their best interests. There is no clear divide between academic and industrial interests. Maybe in something more pure where truths are evident (e.g. pure math). But not something like machine learning, where your success and funding depends on armies of grad students fine-tuning stuff like the number of "hidden units" in some over-complicated model whose ultimate goal is to over fit a training set and cause more hype, etc.
Nothing wrong with this though, progress happens continually, just not linearly: https://en.wikipedia.org/wiki/Hype_cycle
I used to think that the feud over the origins of backprop was an annoying one-off, but it is not. The field has had many examples of lazy citation patterns, including people not actually reading what they cite, and/or ignoring corrective papers like this one.
The example that bothered me most was the "Gaussian processes" boomlet. It played for a couple of years, but it was basically linear prediction theory redone by people who did not adequately cite prior work nor convey its limitations.
What I'm curious about is how such hype cycles affect our understanding of whether this shit actually works.
As far as I can tell, the "great advance" attributed to deep learning is doing benchmarks better than other methods. Which seems to have the result of doing the same old machine learning tasks more accurately than previously.
The main thing we hear about is more "edgy" stuff like learning video game by oneself, describing pictures with phrases or giving answers to "philosophy". But since these applications are easy to cobble together in a half-assed fashion, is there really more here than incremental progress?
I wouldn't know one way or the other, so I'd like to find out.
We are inventing new "real world" benchmarks to try and counteract this (MS COCO, dialog datasets, the big flickr datasets, translation generally), but many approaches from the 90s were clearly right mathematically (as Jurgen says) and just needed more data fuel. So it is obvious to go back and find interesting ideas that didn't get their due as long as proper attribution is given. Profs also have their pet projects that didn't quite pan out, and often want to breathe new life into a cool idea.
These things are only easy to half-assed cobble together in hindsight AND/OR if you have expertise - having the knowledge and know how to input conditional information, interpreting deep networks as modeling joint probability distributions, etc. is just as much algorithmic design as any other task in graphical modeling, statistics etc. Slapping a big deep convnet (or feedforward net) on new datasets IS easy, and usually not interesting scientifically, but also doesn't get published and is reserved for the blogosphere or bad ArXiV papers.
Incremental progress is 0.5% performance gains in major benchmarks like ImageNet etc. - company PR (and university PR as well) will crow about this but no one in academia really cares unless it is accompanied by interesting scientific ideas or fundamental questions being answered.
Not saying he's wrong, just FYI.
DeepMind is a bit of an exception to this. At least one of the founders was involved in quite a bit of original research way back then.
This phenomenon is common in theoretical computer science. Timing and marketing matter a lot when it comes to getting credit for important inventions. I've seen it many times.
A few of the ideas at play here are:
* Wanting to appear more cutting edge than is actually the case
* Limiting or strengthening patent applicability
* Preventing loss of focus via competitors research
I think researchers should actually get penalized for having a deficient bib.I'm telling this anecdote because, even when I agree that we are forgetting to mention a lot of names, that "PR" work that Hinton et all did, was necessary (IMO) to bring ANN back to the mainstream area.
Or maybe not..
"His formal theory of creativity & curiosity & fun explains art, science, music, and humor."
I've also read papers of his that take completely off-the-wall pot-shots at other researchers.
"His formal theory of creativity & curiosity & fun explains art, science, music, and humor."
Maybe he's a mad man and maybe he just hasn't tweaked his "theory of humor" quite enough to know some people won't get it.
Several founders from Deepmind where his PhD students.
Much of Dr. Schmidhuber's work is very interesting and especially relevant now that RNNs are really heating up again - but it is sometimes hard to figure out exactly which of his papers to cite because many are partially relevant. And having a full page of only Schmidhuber citations is no good either...
Speaking as a member of the Montreal lab, I am much more up to date with the work that happens here - so it is hard to fight the natural tendency to cite recent papers you know (since they all came from work you know of, cause you were there). Notice too that all 3 (Hinton, LeCun, and Bengio) worked directly together at some point, and collaborated often beyond that. So a version of this is in effect, whereas Juergen has been more separated (both geographically, and work focus wise) than the other 3. NYU Toronto and Montreal are all in an 8 hour triangle!
Not to take anything away from his points (I try to cite as many of his papers as possible without seeming ridiculous, generally) but these are the general factors at play. We cannot possibly cite every paper in the field, and shining the light on new works can be more important than citing older work AS LONG AS there is no claiming as a pure innovation work that was already done "in the nineties".
Claiming to improve some technique or take it from curious to usable is more than fine - but given the recent deep learning hype even recent papers are getting overshadowed by others claiming some new innovation which already exists in very current literature.
Especially given the work that is coming out of industrial labs (Google, FB, MSR, etc.) it is fairly frequent to see the same model being touted as new (with minor citations if lucky) when the exact same technique first appeared 6 months ago. Being well-read is not an option as an academic - it is a requirement! The PR machine of these companies is unfortunately very effective at dominating the airwaves if you have competing or related work, especially if you are not from a school with good press e.g. MIT, Stanford.
On the one hand, at the conceptual level, unless you are at the cutting edge of CS theory, I'm pretty sure almost anything else that is done in computer science is a mere re-wording of something that was done in the 1970-1980s. So there is no "holier than thou" at this level.
On the other hand, in terms of practical results in context, there are many important consequences of being able to take old concepts and run them faster, because the hardware has improved and well, generally, the entire world is different.
A big part of "popularizing" a technique is having a good implementation that takes advantage of advances in computing speed. So the author of the article misses the practical value of popularizing.
At the same time, what he says is valuable because he touches on a fundamental choice that we all make: do you want to be a groundwork layer or a popularizer?
The problem of course is that groundwork layers are mostly forgotten, with their contributions recognized posthumously, as that's how far out you have to be to lay any new groundwork, and it's difficult to predict what will be the foundation for the next hundreds of years.
It's not just deep learning, it's basically that anything that becomes popular enough to be noticed here probably has a long history behind it, and if we are to move forward we need to be in the headspace of those who had the sense back then to form it, and not be in the space of popularizing or being the tool of the popularizer.
There's no denying that he is a brilliant pioneer, and his accomplishments are indeed incredible, but that does not translate into status which is probably very upsetting.