He already have a Turing award and don't give a rat's ass about who owns how much search traffic. OpenAI just like Google will give him millions of dollar just to be a part of organization
> Explicit, efficient error backpropagation (BP) in arbitrary, discrete, possibly sparsely connected, NN-like networks apparently was first described in a 1970 master's thesis (Linnainmaa, 1970, 1976), albeit without reference to NNs. BP is also known as the reverse mode of automatic differentiation (e.g., Griewank, 2012), where the costs of forward activation spreading essentially equal the costs of backward derivative calculation. See early BP FORTRAN code (Linnainmaa, 1970) and closely related work (Ostrovskii et al., 1971).
> BP was soon explicitly used to minimize cost functions by adapting control parameters (weights) (Dreyfus, 1973). This was followed by some preliminary, NN-specific discussion (Werbos, 1974, section 5.5.1), and a computer program for automatically deriving and implementing BP for any given differentiable system (Speelpenning, 1980).
> To my knowledge, the first NN-specific application of efficient BP as above was described by Werbos (1982). Related work was published several years later (Parker, 1985; LeCun, 1985). When computers had become 10,000 times faster per Dollar and much more accessible than those of 1960-1970, a paper of 1986 significantly contributed to the popularisation of BP for NNs (Rumelhart et al., 1986), experimentally demonstrating the emergence of useful internal representations in hidden layers.
https://people.idsia.ch/~juergen/who-invented-backpropagatio...
Hinton wasn’t the first to use NNs for language models either. That was Bengio.
[1]Learning representations by back-propagating errors
You've now gone from one false claim "he literally invented backpropagation", to another false claim "he is one of the first people to use it for multilayer perceptrons", and will need to revise your claim even further.
I don't particularly blame you specifically, as I said the field of ML is so bad when it comes to properly recognizing the teams of people who made significant contributions to it.
In an alternate universe, NNs are still slow and compute limited, and we use something like evolutionary algorithms for solving hard problems. Hinton would still be just as smart and backpropagation still just as sound but no one would listen to his opinions on the future of AI.
The point is, he is quite lucky in terms of time and place, and giving outsized weight to his opinions on matters not directly related to his work is a fairly clear example of survivorship bias.
Finally, we also shouldn’t ignore the fact that Hinton’s isn’t the only well-credentialed opinion out there. There are other equally if not more esteemed academics with whom Hinton is at odds. Him inventing backpropagation is good enough to get him in the door to that conversation, but doesn’t give him carte blanche authority on the matter.
That is not at all a slam dunk argument. It’s barely anything.
My main point wasn’t to undermine Hinton by saying he was lucky. I did do that and I stand by it. But my main point was to say that to a large degree the future on this issue is unknowable because it depends on so many crucial yet undetermined factors. And there’s nothing you could know about backpropagation, neural networks, or computer science in general which could resolve those questions.
The CTO at OpenAI is https://en.wikipedia.org/wiki/Mira_Murati who does not have a PhD.
I don't think Hinton cared about fame as you imagined.
I never took it but it will be interesting to see what kind of synthesis between traditional logic and neural network paradigms can be achieved.
Ben Goertzel talks about his work on something like this at around the 16 minute mark in this video:
One of GOpenAI's recent breakthroughs was switching to FlashAttention, invented at Stanford and University at Buffalo.
The fact that he never said anything before and the fact that he's saying something now means two things in my mind:
1. He is noticing something different about the current iteration of AI technology. We crossed some threshold.
2. Hinton is being honest.[1]: https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
That being said, architectures are also important when they can reduce computational complexity by orders of magnitude.
Maybe, but there is another force at play here too. It's that journalists want stories about AI, so they look for the most prominent people related to AI. The ones who the readers will recognize, or the ones who have good enough credentials for the journalists to impress upon their editors and readers that these are experts. The ones being asked to share their story might be trying to grab the limelight or be indifferent or even not want to talk so much about it. In any case I argue that journalism has a role. Probably these professional journalists are skilled enough that they could make any average person look like a 'limelight grabber' if the journalist had enough reason to badger that person for a story.
This isn't the case for everyone. Some really are trying to grab the limelight, like some who are really pushing their research agenda or like the professional science popularizers. It's people like Gary Marcus and Wolfram and Harari and Lanier and Steven Pinker and Malcolm Gladwell and Nassim Taleb, as a short list off the top of my head. I'm not sure I would be so quick to put Hinton among that group, but maybe it's true.
Your post suggests that you know almost nothing about how modern deep learning originated.
I’m sure he’s smart and all. His contributions were valuable. But he’s not special in this particular moment.
Who, in your opinion, _does_ have enough context to be worth our attention?
Because if you’re waiting for Sam Altman or the entire OpenAI team to say “guys, I think we made a mistake here” we’re going to be knee-deep in paperclips.
https://en.wikipedia.org/wiki/Thermonuclear_weapon
https://www.simonandschuster.com/books/Dark-Sun/Richard-Rhod...
Today:
Germany's relevant minister has already declared at G7 that Germany can not follow Italy's example. "Banning generative AI is not an option".
https://asia.nikkei.com/Spotlight/G-7-in-Japan/Banning-gener...
US Senate has a bill drawing the line on AI launching nuclear weapons but to think US military, intelligence, and industry will sit out the AI arms race is not realistic.
https://www.markey.senate.gov/imo/media/doc/block_nuclear_la...
China's CPC's future existence (imo) depends on AI based surveillance, propaganda, and realtime behavior conditioning. (re RT conditioning: We've already experienced this outselves via interacting with the recent chatbots to some extent. I certainly modulated my interactions to avoid the AI mommy retors.)
https://d3.harvard.edu/platform-rctom/submission/using-machi...
But beyond all of that, what are we really asking? Are we asking about social ramifications? Because I don’t think the OpenAI devs are particularly noteworthy in their ability to divine those either. It’s more of a business question if anything. Are we talking about where the tech goes next? Because then it’s probably the devs or at least indie folks playing with the models themselves.
None of that means Hinton’s opinions are wrong. Form your own opinions. Don’t delegate your thinking.
Are you basically saying that you only trust warnings about AI from people who have pushed the most recent update to the latest headline-grabbing AI system at the latest AI darling unicorn? If so, aren't those people strongly self-selected to be optimistic about AI's impacts, else they might not be so keen on actively building it? And that's even setting aside they would also be financially incentivized against publicly expressing whatever doubts they do hold.
Isn't this is kind of like asking for authoritative opinions on carbon emissions from the people who are actually pumping the oil?
If not Geoff Hinton, then, who?
Ultimately the harm is either real or not. If it is real, then the people with the most accurate beliefs and principles will be the ones who never joined the industry in the first place because they anticipated where it would lead, and didn't want to contribute. If it is not real, then the people with the most accurate beliefs will be the ones leading the charge to accelerate the industry. But neither group's opinions carry much credibility as opinions, because it's obvious in advance what opinions each group would self-select to have. (So they can only hope to persuade by offering logical arguments and data, not by the weight of their authoritative opinions.)
In my view, someone who makes landmark contributions to the oil industry for 20 years and then quits in order to speak frankly about their concerns with the societal impacts of their industry... is probably the most credible voice you could ever expect to find expressing a concern, if your measure of credibility involves experience pumping oil.
He received a Turing Award for his work that was foundational to the current state of the art.
Isn't that ad hoc ergo propter hoc?
That argument would also support the statement "he went all in with 2-7 preflop, and won the hand, so he must be good at poker" -- I assume you and I would both agree that statement is not true. So why does it apply in Geoffrey's case?
With respect, you seem to be shifting goalposts, from the indefensible (Hinton doesn't know what he's talking about) to the irrelevant (Hinton doesn't have perfect and complete knowledge of the future).
Who knows, maybe a decade or two from now we'll see a resurgence of capsule networks, or maybe not. But I'd be a bit more careful about rejecting Hinton's hunches out of hand, his track record is pretty good.
Even if they're too busy doing the work, they're still thinking about what it would be like if it performed successfully, and it does seem to always take more retrospection before a leader can fully raise their head and more carefully consider unintended consequences.
Early success can give the impression that future efforts have difficulty being as meaningful, but also realistically after that the successful individual does not need to struggle to prove themself any more the way the less-accomplished would be expected to do.
Then there's seniority itself, and maturity levels that can not be gained any other way.
Beyond that when retirement is within easy reach you don't really have the same obligation to decorum itself as you would earlier, in order to actually maintain the same desired level of decorum.
Dr. Hinton seems to do a pretty good job of comparing himself to Oppenheimer.
I don't see how anyone else can question his standing more seriously than that.
Schmidhuber claims to have invented something formally equivalent to the linear Transformer architecture (slightly weaker) years before:
It is still winner takes all, but if you look at the overall landscape, there are plenty of opportunities where you can have an outsized impact - you can have localized fame and fortune (anyone with AI expertise under their belt have no problems with fortune!)
Example: https://www.microsoft.com/en-us/research/wp-content/uploads/... (Michele Banko published a few similar papers on that topic)
The success of these LLMs comes down to the Transformer architecture which was a bit of an accidental discovery - designed for sequence-to-sequence (e.g. machine translation) NLP use by a group of Google researchers (almost all of who have since left and started their own companies).
The "Attention is all you need" Transformer seq-2-seq paper, while very significant, was an evolution of other seq-2-seq approaches such as Ilya Sutskever's "Sequence to Sequence Learning with Neural Networks". Sutskever is of course one of the OpenAI co-founders and chief scientist. He was also one of Geoff Hinton's students who worked on the AlexNet DNN that won the 2012 ImageNet competition, really kicking off the modern DNN revolution.
He claims to have ideas for architectures that could surpass the capabilities of GPT4, but can't try them for a lack of funding in his academic setting. He said his ideas were nothing short of genius..
(unfortunately german) source: https://science.orf.at/stories/3218956/
Hinton has had a significant part to play in the current state of the art.
The problem has always been, and now will likely always be, the hardware. I've written about this at length in my previous comments, but a split happened in the mid-late 1990s with the arrival of video cards like the Voodoo that set alternative computation like AI back decades.
At the time, GPUs sounded like a great way to bypass the stagnation of CPUs and memory busses which ran at pathetic speeds like 33 MHz. And even today, GPUs can be thousands of times faster than CPUs. The tradeoff is their lack of general-purpose programmability and how the user is forced to deal with manually moving buffers in and out of GPU memory space. For those reasons alone, I'm out.
What we really needed was something like the 3D chip from the Terminator II movie, where a large array of simple CPUs (possibly even lacking a cache) perform ordinary desktop computing with local memories connected into something like a single large content-addressable memory.
Yes those can be tricky to program, but modern Lisp and Haskell-style functional languages and even bare-hands languages like Rust that enforce manual memory management can do it. And Docker takes away much of the complexity of orchestrating distributed processes.
Anyway, what's going to happen now is that companies will pour billions (trillions?) of dollars into dedicated AI processors that use stuff like TensorFlow to run neural nets. Which is fine. But nobody will make the general-purpose transputers and MIMD (multiple instruction multiple data) under-$1000 chips like I've talked about. Had that architecture kept up with Moore's law, 1000 core chips would have been standard in 2010, and we'd have chips approaching 1 million cores today. Then children using toy languages would be able to try alternatives like genetic algorithms, simulated annealing, etc etc etc with one-liners and explore new models of computation. Sadly, my belief now is that will never happen.
But hey, I'm always wrong about everything. RISC-V might be able to do it, and a few others. And we're coming out of the proprietary/privatization malaise of the last 20-40 years since the pandemic revealed just how fragile our system of colonial-exploitation-powered supply chains really is. A little democratization of AI on commoditized GPUs could spur these older/simpler designs that were suppressed to protect the profits of today's major players. So new developments more than 5-10 years out can't be predicted anymore, which is a really good thing. I haven't felt this inspired by not knowing what's going to happen since the Dot Bomb when I lost that feeling.
>... Docker takes away much of the complexity of orchestrating distributed processes.
The T-800 running on Docker: After failing to balance its minigun, it falls forward out of the office window, pancaking in the parking lot below. Roll credits.