Yann LeCun, Geoffrey Hinton and Yoshua Bengio win Turing Award
nytimes.com
nytimes.com
You could instantly see the results he presents were way better than what was state of the art at that time. Amazing.
---
[0] Grant Sanderson (3Blue1Brown) started a youtube-series covering Neural Networks (4 episodes, so far) that helps gain an intuitive grasp on the topic: https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700...
It's the same with Hinton et. all. They were among many other pioneers, but what really sets them apart is that they made it all work. They also analyzed why it works.
Seems disingenuous to not mention that this device only worked due to ground effects and "flew" 8 inches off the ground (according to Wikipedia).
Werbos and others were very aware that the method worked and also of the implications of it:
> This approach makes it possible to develop generalized, adaptive artificial intelligence, capable of achieving results comparable to what is discussed in science fiction
http://werbos.com/Neural/SensitivityIFIPSeptember1981.pdf
The Hinton cartel happened to be in the right condition that allowed them to ignore the politically correct BS starting at the turn of the century which rendered the study of neurons, IQ and intelligence massively unpopular; and ruthless enough to only cite one another and call themselves the "fathers of deep learning" rather than actually citing the originators of these ideas.
https://en.m.wikipedia.org/wiki/Gustave_Whitehead
Ive seen some debate around Gustave Whitehead. There apparently is some evidence that he achieved powered flight as early as 1899.
It is claimed, he did not speak English very well. So it’s suspected he just didn’t publicize his work as much.
Replicas of his early flying machine designs have been shown to fly. Including one from 1901:
People deserve credit for taking an idea and running with it. Without Hinton's efforts (in 1987, 2006, or 2012), deep learning would have never seen the explosive growth.
Stating "would never have" is nothing more than telling a story.
Yes, as popularizers.
> deep learning would have never seen the explosive growth
This is extremely unlikely to be true. I think the explosive growth occured because neurons and intelligence were unpopular in the 2000s because of PC, deferring progress in this area to occur explosively later. The Hinton cartel were not the only people who were aware that AI research "has become silly" (in his own words).
That said, I don't know how you could argue against Hinton's impact on the field, or LeCun and Bengio's achievements.
Unfortunately, a good PR is definitely important for prizes.
1) Some scrappy researcher on a low budget develops a new AI technology that shows promise in a specific area.
2) More researchers take that idea and successfully apply it more broadly.
3) Massive investment in research happens, pushing PhD candidates into extracting every possible nuance out of that technology.
4) Researchers, encouraged by the broad success, begin to think we've found the one true key in the quest towards general intelligence. Suddenly all the researchers are Neats. The research funding is now mostly businesses instead of governments and non-profits. They grow overconfident that it can be used for virtually anything, funding everything in sight.
5) Starry-eyed futurists start lionizing the founders of the new revolution as geniuses, publishing flowery bullshit about how the world will be forever changed and how the General AI revolution is only 5-10 years away.
6) Surprise! It actually can't be used that broadly. Improvements stall, and CEOs start to realize that they can't just dump data in and get money out. It becomes a commercial disappointment, triggering disinvestment.
7) The Neats, disappointed that their glamorous and simplistic theory of general intelligence, start to disappear. Newspapers begin to make fun of all of the visionaries, comparing them to the people that said flying cars were 5-10 years away.
8) A handful of Scruffies take over the now low-funding AI winter, working hard on minimal budgets until a new breakthrough is found and we return back to #1.
This article is showing that we're solidly within stage 5, and we're already seeing signs of stage 6. When everybody is buying, it's time to sell.
Part of that too, is that I think people now realize that narrow AI is sufficient to create tremendous value, and that it doesn't necessarily matter if the "AGI breakthrough" happens anytime soon or not.
It's hard to be sure, but I don't see the kind of collapse that happened in the past happening anytime soon.
We haven't stopped buying tulips, trains, stocks, technology, or real estate. And just as well, we haven't stopped using symbolic AI, single-layer neural nets, ensemble models, expert systems, or logic programming languages. The new topological enhancements to Neural Nets won't ever go away either...but that doesn't mean we won't see a drop in investment once the general public realizes that your neural nets aren't going to learn how to do double-entry accounting any time soon. The AI winter isn't characterized by the technology going away, it is characterized by lofty idealism being shattered and investment dropping back to reflect reality.
Right, but that doesn't match the way I feel the term "AI Winter" has been used. Now I could be mis-interpreting things, but I've always looked at an "AI Winter" as a period of underinvestment, created as an over-reaction to the mechanics you're referring to.
but that doesn't mean we won't see a drop in investment once the general public realizes that your neural nets aren't going to learn how to do double-entry accounting any time soon.
Right, but again, I don't think most people use "AI Winter" to mean a simple "drop in investment". If it were something that straightforward, there would be no need for the "Winter" metaphor.
And that's why I say there may indeed be a drop... an "AI Fall" if you will, that still represents a pull-back of sorts, but perhaps just a less pronounced and extended pullback like we've seen in the past.
less pronounced and extended pullback like we've seen in the past.
should read:
less pronounced and extended pullback than we've seen in the past.
Now it feels to me we are in a design winter again: there are still a lot of designers on staff, but they are relegated to a "touch up" role... They have been moved ahead of implementation (which is good) but they are expected to "touch up" concepts generated by management in a matter of days. Very similar to when designers would get a few days to clean up a UI that was already implemented.
Sort of off topic, but just and example that widespread use and business value doesn't mean something can't fall back into a winter of sorts.
>On the other hand, if you bet that stocks will go down, you have some compensating psychic rewards. For one thing, occasionally stocks will go down, and you will be praised for your prescience in predicting the crash, and the people who were long will be mocked for their complacency. How smart you will feel!
>For another thing, even if stocks haven’t gone down, you get to borrow psychically, as it were, against that future moment of glory. You can just go around sort of saying “this is unsustainable and eventually stocks will go down and I will be praised for my prescience,” and people will be surprisingly willing to say “yes that’s correct, I admire your hypothetical prescience.” Particularly since the 2008 financial crisis, financial markets—and financial media—have a strongly entrenched narrative of prescient bears and complacent bulls, a widespread sense that any rising market is suspect and that the cynical view is always the smart one.
It's just weird to me there's always people looking to figure out a way to call the top on absolutely everything.
NNets are already a disappointment to nearly any applied researcher that isn't sitting on petabytes of data. It won't be long before the CEOs realize it and put their research funding somewhere else.
Who are "people" in this context? Are you talking about the general public, or misinformed journalists who are writing about things they don't understand? Because from what I've seen, by and large, the researchers working on Deep Learning are not claiming that DL (alone) is sufficient to achieve AGI, nor do they posit that AGI is anywhere close to reality. Read, for example, Martin Ford's book Architects of Intelligence[1] and note how the various researchers interviewed talk about AGI. That list, BTW, includes all three of LeCun, Hinton, and Bengio, as well as many others.
Now if you talk to the actual researchers/engineers when they're not afraid of hiding behind their NDAs (basically a friend at a bar), these people will talk about the challenges and how much work there is to do and so on. But the company's public stance, and the many billions of dollars chasing this area, is that it's just around the corner.
The point is, there is a lot of hype and it's easy to understand why. Elon Musk is happy to proclaim that Tesla is already selling the hardware for full self-driving vehicles, and it's just the regulators holding him back (which is absurd, to put it kindly). This is the same person saying AI is a bigger threat than nuclear weapons. So arguably the largest tech celebrity isn't even talking about when, he's saying it's here and we need to be prepared. This is the kind of person shaping the public's opinion.
But the money chasing these investments are mostly from private VC firms and tech companies, not public research grants. Once investors realize AI just isn't good enough for self-driving cars right now they won't want to keep throwing billions of dollars behind it. And that's when we arrive at a winter. Even Alphabet is starting to get weird about raising its stock price, so not even the tech giants are immune from this.
There are many other verticals that follow the same story, I was just using the most obvious example of an industry that has billions flowing into it despite nobody having a real product on the promise of advancements in AI and "the time is now!" But eventually (my guess is next year, but I'm old enough to know it's impossible to time these things) the hype will come crashing down to reality and that is when funding will dry up.
I was at a local government plenary session at Chengdu (long story) a couple of year back. The lead speaker kept waxing about how China invented and contributed AI to the world.
I mentioned some points about about China's AI contributions/influence :
0. Algorithms -> The OGs as celebrated by this Turing Award (Canadian, but mostly led by American universities. On thing I did mention was due to the sheer factory production of Chinese PhDs, there is a lot of stuff arxiv from China )
1. Frameworks -> Tensorflow, PyTorch (American)
2. Hardware -> Nvidia (American)
3. Distribution -> Github (American)
4. Education -> Medium, Github, Youtube, Fast.ai (American)
5. Cloud -> AWS, Google, Azure (American)
China has some equivalent for all of them, ie. PaddlePaddle by Baidu, AliCloud, etc. but none of them have the reach, influence or domination of the American counterparts.
Suffice to say that my points weren't taken too well and have been disinvited since then.
China has done some fantastic work in ML, but I think the award was a long time coming for these three.
Me and my friends who also work in this field are generally excited about ML/AI getting more recognition from the CS community at large btw.
LSTM's are clever and they were bleeding edge up to 2014, but once they were understood better as a bypass mechanism, attention, context vectors and averaging networks and causal convolution are starting to replace them.
As far as I'm aware, causal convolutions were used in WaveNet (and subsequent models) and a small number of NLP applications. Meanwhile, LSTM-based models are used in just about every NLP paper, and at least a baseline in the newer ones more dominated by Transformers.
Out of curiosity, at the same level as any of the other three? I could see it from Hinton, maybe LeCunn, but Bengio?
But that's exactly the point of the Transformer model, with a paper aptly titled "Attention is all you need" [1]. And the Bert architecture, based in this idea, seems to be doing well. And they claim to be bery flexible, too[2].
Maybe that's what you meant with "unless you brutely search over hundreds of hyperparameters configs", but then again, isn't that what NNs are about anyway?
"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
When you put researchers or groups into order of importance, the work from this trio comes before others, including Schmidhuber. The Turing Award is given to those at the top and it's not inclusive.
If I could give Turing Award for someone in the field who has not received it yet, I would give it to Vladimir Vapnik.
So would I, but I don't think that will happen. He's is pretty much an antithesis to everything Deep Learning is about. Theory-first over trial-and-error, math over intuition, small datasets over big data, advances in understanding vs advances in results.
I haven't seen a single discussion of his paper on combining classifiers[1] anywhere on the web.
[1] - http://jmlr.csail.mit.edu/papers/volume17/16-137/16-137.pdf
Hinton/Lecun pioneered early BP research, which is significant and fundamental. Bengio has attention mechanism/GAN, and in early days greedy layer wise pretraining for RBM (which isn't something end up today, but pertaining is a thing back in pre 2010 era).
Not saying LSTM isn't significant, but his achievements aren't quite there with those 3, looking closer.
http://people.idsia.ch/~juergen/deep-learning-conspiracy.htm...
Note for people put off by the link: The conspiracy refers to the title of a paper by the Turing award winners, not Schmidhuber's.
That's not true with Schmidhuber.
Imagenet challenge and 2012 has been an inflection point but few have slogged through the AI winter and made it through to the other side.
> On 30 September 2012, a convolutional neural network (CNN) called AlexNet achieved a top-5 error of 15.3% in the ImageNet 2012 Challenge, more than 10.8 percentage points lower than that of the runner up. This was made feasible due to the utilization of Graphics processing units (GPUs) during training, an essential ingredient of the deep learning revolution. According to The Economist, "Suddenly people started to pay attention, not just within the AI community but across the technology industry as a whole."
Also, there is educational value in conversation.
What is the value to you in shutting down a conversation someone else is interested in having?
At least cite the answer to the question if you feel it's settled scholarship.
If you wanted a first-rate CS education, you could do a lot worse than to go through the winners of the Turing Award[0] and review their seminal works. I've only sampled maybe half of them but every time I pick one I learn something interesting.
Curious: Apart from Jurgen Schmidhuber, who else do you think were left out and why do you think so?
Text, language, human chat
Image recognition from blurry, multiple views, multiple lighting conditions photos
formal patterns with large numbers of variations
These are not at all the same, yet the praise seems to want to declare "the best" and "beating competitors" .. why is something so multi-faceted, reduced to the logic of a sports event ?
we must glorify the efforts of a arbitrarily chosen single individual, in the hope of someday becoming that lauded singleton, the one allowed by decree to piss on the heads of those fools below who didn't win the race.
It's pretty much the same in ML. These people built the building blocks that we still use in ML everyday (e.g. backpropagation, ConvNets etc). But the fine-tuning, packaging and tooling of these techniques also changes all the time and it can be hard to keep up. Having moved from full-stack web to ML I have a similar feeling about the pace of things.
However, these are 3 people who overcame a lot of nay-sayers to prove some something could actually work. I remember back in university the lecturer said that neural networks couldn't scale because of the vanishing gradient problem.
It was actually very easy to be critical, you want fit a massive number of parameters which are not-very-orthoganal to an under constrained problem? Sounds pretty dumb to me!
I think these 3 deserve recognition for their tenacity and the new world of "under defined gradient decent", sometimes called DL, which they opened up.
As it should be. The whole AI buzz misleads people into thinking we are close to solving things we actually are not. Like upload filters capable of differentiating fair use and copyright infringement.
[0]: https://amturing.acm.org/award_winners/hinton_4791679.cfm
It's not.
It's about some incremental progress in making neuronal networks classify stuff.
Google trends seems to indicate the opposite.
https://trends.google.com/trends/explore?q=turing%20award,tu...
But it's not about "well known" anyhow. It's about significance. A turing test winning chat bot would be a much bigger breakthrough then an incremental improvement in image and text classification.
It could mean no one needs to Google turing award because everyone knows what it is and don't know what the Turing test is.
Nah.
It was the high performance scientific computing community that started to use graphics cards to perform matrix computations in the early 2000's. The term GPGPU term was invented. GPU programming pain in the ass in early 2000's. Nvidia saw some markets in GPGPU's for scientific computing. First release of CUDA was 2007.
NVidia didn't do science? It depends on what you call science (math and computer science are not "science" according to a very strict definition involving falsifiability). But in any case, an enormous amount of research went into the development of GPUs. It's not just engineering work. Lots of PhD (students) contributed.
Hardware served as a catalyst, but it was not a necessity.