This is a different situation. There's the "NeurIPS 2024 Test of Time Paper Awards" where they award a historical paper. In this case, a paper from 2014 was awarded and his talk is about that and why it passed the test of time.
https://blog.neurips.cc/2024/11/27/announcing-the-neurips-20...
The title chosen for the HN submission leaves out that important context. So that's why you are disappointed now.
Obviously Perceptrons came out well before 2003, but I don't think it's necessarily out of line to say that they had limited efficacy before then, both for theoretical and compute reasons. But maybe I'm misunderstanding your criticism?
There were precursors. At least Ehud Shapiro's doctoral thesis ("Automated Debugging") in the 1980's and Gordon Plotkin's doctoral thesis in the 1970's ("Automated Methods of Inductive Inference"). Sorry for not giving the exact years off the top of my head but I think it was 1983 and 1976, respectively.
The point you are making is very right however because modern machine learning as a field started in the 1980's with the fall of expert systems, in fact it basically started as an effort to overcome one of the major limitations of expert systems, the so-called "knowledge acquisition bottleneck", which is to say, the difficulty of creating and maintaining huge databases of expert knowledge (in the form of production rules).
In any case the seminal textbook in the field for the first 20 years, Tom Mitchell's Machine Learning came out in 1997 (https://www.cse.iitb.ac.in/~cs725/notes/slides/tom_mitchell/...) and includes probabilistic, neural-net based and symbolic, logic-based approaches. So not only machines could "learn" way before 2003 but they could also learn in many different ways than what Ilya Sutskever means.
We can go further back, to Donald Michie's 1961 MENACE (the first Reinforcement Learning system, implemented on a computer made of matchboxes with coloured beads used to encode state) and Arthur Samuel's 1959 checkers player (a paper on which gave the name to the field of machine learning).
Lots of learning all over the place long, long before 2003.
Probably what we are discussing here is not the next breakthrough...
* Before the current renaissance of neural networks (pre ~2014ish), it was unclear that scaling would work. That is, simple algorithms on lots of data. The last decade has pretty much addressed that critique and it's clear that scaling does work to a large extent, and spectacularly so.
* Much of the current neural network models and research are geared towards "one-shot" algorithms, doing pattern matching and giving an immediate result. Contrast this with search which needs to do inference time compute or search.
* The exponential increase in power means that neural network models are quickly sponging up as much data as they can find and we're quickly running into the limits of science, art and other data that humans have created in the last 5k years or so.
* Sutskever points out, as an analogy, nature has created a better model for humans (the brain to mass ratio for animals) with hominids finding more efficient compute than other animals, even ones with much larger brains and neuron count.
* Sutskever is advocating for better models, presumably focusing on inference time computer more.
In some sense, we're coming a bit full circle where people who were advocating for pure scaling (simple algorithms + lots of data) for learning are now advocating for better algorithms, presumably with a focus on inference time compute (read: search).
I agree that it's a little opaque, especially for people who haven't been paying attention to past and current research, but this message seems pretty clear to me.
Noam Brown had a talk recently titled "Parables on the Power of Planning in AI" [0] which addresses this point more head on.
I will also point out that the scaling hypothesis is closely related to "The Bitter Lesson" by Rich Sutton [1]. Most people focus on the "learning" aspect of scaling but "The Bitter Lesson" very clearly articulates learning and search as the methods most amenable to compute. From Sutton:
"""
...
Search and learning are the two most important classes of techniques for utilizing massive amounts of computation in AI research.
...
"""
[0] https://youtube.com/watch?v=eaAonE58sLU
[1] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
"We've made a copy of the internet, run current state of the art methods on it and GPT-O1 is the best we can do. We need better (inference/search) algorithms to make progress"