Geoffrey Hinton spent 30 years on an idea many other scientists dismissed
torontolife.com
torontolife.com
This is so disrespectful to the 1000s of researchers who have been studying machine learning since well before 2012. It was well established that the future of teaching computers came from statistics and not rules in the 80s/90s by researchers like Michael Jordan (https://en.wikipedia.org/wiki/Michael_I._Jordan) and his students.
It was even engrained in popular culture: neural networks are how the AI brain worked in Terminator 2, in 1991! https://www.youtube.com/watch?v=xcgVztdMrX4
edit: I don't want to downplay Hinton's accomplishments, I've been lucky to have been surrounded by and motivated by his work since I started learning machine learning. I did my masters research on neural networks that were partly inspired by his work, and it was a deep networks paper he presented at a NIPS 2006 workshop that got me really excited to stay in machine learning while I was starting my career.
I remember that in the early 90s, neural nets were supposed to be the huge new thing. Many scanners shipped with built-in OCR, the Apple Newton had handwriting recognition on a PDA, and my Centris 660AV could do speech recognition & text-to-speech out of the box.
But they ultimately weren't powerful enough to satisfy customers or meaningfully change how people interacted with computers, so they failed in the market, and the hype cycle moved on to the World Wide Web.
I guess this shows the power of continuing to study something when the fad goes away, so that you're well positioned to capitalize when the next fad hits.
So true. VR popped up in the 90/00s (Virtual Boy, VRML, etc.), went nowhere, and here we are again.
I suspect something similar will happen with blockchain/cryptocurrency. Only once all the hype and speculation dies off will meaningful uses for the tech become evident.
Er, logic inference is "only so powerful"? The first order predicate calculus (first order logic) is Turing-complete and there is plenty of maths that prove the soundess and completeness of various logic inference rules. In other words, if you can compute a function, you can compute it with first order logic.
The switch from symbolic to statistical AI happened partly because of the AI winter that cut the funding to all AI, which at the time was primarily symbolic AI, partly because it became evident that developing and maintaining huge databases of logic rules was inefficient. Some of the early work in machine learning focused on overcoming this inefficiency by inducing rules from data.
Brainfuck is turing complete, does that make it as powerful as a general purpose programming language such as Java?
I'm not criticizing first order logic, it has use cases no doubt. On the other hand there are many cases where codifying knowledge into a we formed KB isn't practical. Some of these use cases are better suited to probabilistic approaches, deep learning, and some are unsolved. That pretty clearly puts limits on the power of FOL.
Also, thanks for mansplaining turing completeness.
Snark is not OK:
Be civil. Don't say things you wouldn't say face-to-face. Don't be snarky. Comments should get more civil and substantive, not less, as a topic gets more divisive.
As a counterpoint, one of the most successful (classes of) machine learning algorithms are Decision Tree learners, whose models are decidedly symbolic.
There's also plenty of work on learning first-order logic theories with neural nets (see for instance the work of Artur D'Avilla Garcez).
The problem with rules is that it's hard to develop and maintain large rule-bases. However, it's perfectly possible to do machine learning for rules -and logic-based machine learning is totally a thing (full disclosure: it's the thing I'm doing a PhD on). You probably haven't heard of it though because data scientists tend to be good with statistics but bad with symbolic logic.
AI has had a lot of fads, but neural networks are here to stay, in my opinion.
I have just finished reading this interesting wiki titled AI winter defined as a period of reduced funding and interest in AI research:
https://en.m.wikipedia.org/wiki/AI_winter
There have been such winters since the 60's.
I think that is what the article may be getting it.
https://en.wikipedia.org/wiki/Perceptrons_(book)
Neural nets fell out of favor in the 1970's but came back and became hot in the early 1980's with work by John Hopfield and others that addressed the objections.
https://en.wikipedia.org/wiki/John_Hopfield
Practical and commercial successes were limited in the 1980's and 1990's which led to a reasonable decline in interest in the method. There were some commercial successes such as HNC Software which used neural nets for credit scoring and was acquired by Fair Isaac Corporation (FICO).
https://en.wikipedia.org/wiki/Robert_Hecht-Nielsen
I turned down a job offer from HNC in late 1992 and neural nets were still clearly hot at that time.
Some people continued to use neural nets with some limited success in the late 1990's and 2000s. I saw some successes using neural nets to locate faces in images, for example. Mostly they failed.
AI research is very faddish with periods of extreme optimism about a technique followed by disillusionment. One may wonder how much of the current Machine Learning/Deep Learning hype will prove exaggerated.
Also, traditional Hidden Markov Model (HMM) speech recognition is not rule based at all. It uses a maximum likelihood based extremely complex statistical model of speech.
https://www.semanticscholar.org/author/Geoffrey-E-Hinton/169...
Citations to his papers have been rising steadily from between 88 and 107 in 1987 to between 685 and 826 in 1999. That's hardly an unpopular researcher.
And for a bit of comparison with other machine learning researchers, here's a link to a data set of family relations from a 1986 paper by Hinton:
https://archive.ics.uci.edu/ml/datasets/Kinship
At the bottom of that page, in the Relevant Papers sections there's two links to two papers using the data set, one Hinton's own paper that introduces it and one by Quinlan.
Clicking on the [Web Link] links for the two papers, I can see the references to those papers. There is a single reference to Quinlan's paper. There are 43 to Hinton's, of which all but 6 are from 1999 and earlier. And those are not self-references, neither references by Bengio, Le Cun et al. If there is a clique, it is hard to see it.
So there was a lot of interest to Hinton's work even in the years he was supposed to be "exiled to the academic hinterland" as another article said.
Truth! I happened to go through grad school at a time when SVMs and kernel methods were cool, neural nets were the opposite, and Hinton and his students were oddball outcasts at NIPS. Kudos to Hinton for sticking to his story until the current "deep learning" hype wave presumably allowed him to buy a yacht and an island. I imagine we'll be hearing about some other non-linear optimization technique in a few years.
"What the book does prove is that in three-layered feed-forward perceptrons (with a so-called "hidden" or "intermediary" layer), it is not possible to compute some predicates unless at least one of the neurons in the first layer of neurons (the "intermediary" layer) is connected with a non-null weight to each and every input. This was contrary to a hope held by some researchers in relying mostly on networks with a few layers of "local" neurons, each one connected only to a small number of inputs. A feed-forward machine with "local" neurons is much easier to build and use than a larger, fully connected neural network, so researchers at the time concentrated on these instead of on more complicated models."
Neural Networks weren't really thought of very highly until CNNs started winning image recognition competitions in the early 2010's.
I think most people had the feeling that they were interesting tools to learn how the brain worked, but too slow and opaque to be practical statistical tools. I've been following Hinton for a while (because of Hinton and Shallice 1991), and my understanding is that it was really hard for him to get funding especially when he was just starting out.
The fact that so much of the work from the mid-80's to 2000 came from just a few labs should tell you how hard it was to get funding for that kind of research.
If you look through papers from that era I don't think you'll find that's true at all. You could fairly say that only a few of the prominent labs from the first neural-net boom lasted long enough to still be prominent labs now in the second neural-net boom (though even then there are several: Hinton, Bengio, Schmidhuber, LeCun, etc.). Since they're still around doing interviews and putting out new papers, understandably their work has a higher profile now than that of people who aren't in the field anymore. But there have been a ton of others over the years too, just many of them from the first wave have moved on or retired by now.
Especially in the '90s the field was hot and reasonably large (and it was pretty easy to get funding, too). If you look at e.g. the NIPS 1992 proceedings, it's definitely not just a handful of labs: https://papers.nips.cc/book/advances-in-neural-information-p...
I agree, I don't mean to imply that Hinton did all of the work on NN's, but he certainly contributed a lot to the field, even as the popularity waxed and waned.
Even if you just click on a few of the names in your link, you can see most of those authors submitted ~1-6 papers. Hinton, Bengio, and Jordan have submitted >50 each. I wouldn't say that NN's were exactly thought of as a joke before 2011, but from the people I talked to about them, they weren't considered very promising as practical statistical systems.
When I did my M.S. reasearch in 2010 SVMs were far more popular than neural networks. Deep learning was just over the horizon and most people had given up on it being useful for learning real world problems. We had proven that NNs could learn anything given enough data and the right structure, but the hard part in real life was having enough data and finding the correct structure. So SVMs seemed more robust with available datasets, and most people used those. There were/are several other popular methods, like nearest neighbors and tree-based methods, and NNs just weren't as sexy as they seemed in the late 80's.
It’s not a super long book and an excellent, eye opening read even 50 years later.
If you took half the apps built today and tried to run them on compatible hardware from 20 years ago, they would all flop.
Without the understanding of modern computing, one would have dismissed all such inventions as useless wastes of time.
Probably they'll never mention Friedman and Breiman, which seems pretty unfair considering their gizmos have arguably had a bigger impact in "actual machine learning gizmos deployed..."
But that still skirts the big issue, which is generalization. We are moving, it seems, to transfer learning. The danger is that we don't seem to have a good theory to why it works. At a practitioner level, I don't think this is as much of a problem. For the research, though, it is pretty shaky.
I think there is more than a strong chance this remains the future for a while. And I am layman in this field. At best. But this story presupposes that the past was wrong for not being like the present. That is a tough bar.
Historically, almost all ML approaches were based on separation in Euclidean (vector) space, which is understandable because they were developed for much weaker computers. However, really useful ML tasks require to deal with huge nonlinearities, and the fact that you're training in linear space becomes less relevant.
Neural networks have surmounted the nonlinear difficulty by increasing the number of layers. But it's a question whether a similar result couldn't be achieved with Bayesian networks on binary representations (which is approach I favor).
There is some evidence that the precision of linear calculation in neuron (for example, resolution of the weights) doesn't really make much difference in the performance of neural network. Could it be that the neural learning through vector manipulation is just an artifact of the origin of the neural networks, and the really important thing is the overall organization of the network (the layers)?
Be clear on that. He was not able to see the success using these techniques on older hardware. Period. People weren't dismissing him. They were doing better.
what's with everyone here?
Is that true?
Geoffrey Everest Hinton → Howard Everest Hinton → George Hinton → Mary Ellen Hinton → George Boole
It certainly seems plausible: https://en.wikipedia.org/wiki/George_Boole
"an outsider clinging to a simple proposition: that computers could think like humans do—using intuition rather than rules. "
I stopped reading right there.So why isn't this something that's in the realm of possibility for computers?
At a fundamental level - are our brains actually comparable to how ML works (beyond some basic analogies)? Do we have an statistical engine running inside our heads, needing tremendous "CPU power" to do something remotely useful/accurate?
I'd say that no, and that that conceptual mismatch indicates that the next big iteration on AI will be something more like what D. Hofstadter advocates/researched.
(Using ML as a sidekick, why not. No need to trash out the current progress)