Ilya Sutskever NeurIPS talk [video]
youtube.com
youtube.com
The synopsis, as far as my tired brain can remember:
- Here's a brief summary of the last 10 years
- We're reaching the limit of our scaling laws, because we've trained on all the data we have available on the limit
- Some things that may be next are "agents", "synthetic data", and improving compute
- Some "ANNs are like biological NNs" rehash that would feel questionable if there was a thesis (which there wasn't? something about how body mass vs. brain mass are positively correlated?)
- 3 questions, the first was something about "hallucinations" and whether a model be able to understand if it is hallucinating? Then something that involved cryptocurrencies, and then a _slightly_ interesting question about multi-hop reasoning
I notice with Ilya he wants to talk about these out there speculative topics but defends himself with statements like “I’m not saying when or how just that it will happen” which makes his arguments impossible to address. Stuff like this openly invites the crazies to to interact with him, as seen with the cryptocurrency question at the end.
Right before this was a talk reviewing the impact of GANs that stayed on topic for the conference session throughout.
> the audience is at least partially composed of people with little technical background or AI bros.
I have never seen the term "AI bros". What does it mean?When? It's at an all time high right now
Not recent as in the last 3 months.
I.e. someone who has shallow technical understanding or independent thought, but is following trends in hopes of turning a profit.
NeurIPS is "ruined" by the money and thus attracts huge amounts of people who are all trying to get rich. It's a bloody academic conference people!
Is that available online?
That's just one sentence, but it's pretty important. And while many people already know this, it's important to hear Sutskever say this. So people know it's a common knowledge.
The rest is basically intro/outro.
I guess you're interpreting "self-awareness" in some mythical way, like a soul. But in a trivial sense, they are. Perhaps not to same extent as humans: models do not experience time in a continuous way. But given that it can maintain a dialogue (voice mode, etc), it seems to be phenomenologically equivalent.
It might seem as a complete black box, but we can get some information about it by observing human interactions.
E.g. suppose Алиса does not know English, but she has a bunch of cards with instructions "Turn left", "Bring me an apple", etc. If she shows these cards to Bob and Bob wants to help her, Bob can carry out instructions in a card. If they play this game, the meaning which card induces in Bob's head will be understood by Алиса, thus she will be able to map these cards to meaning in her head.
So there's a way to map meaning which is mediated by language.
Now from math perspective, if we are able to estimate semantic similarity between utterances we might be able to embed them into a latent "semantic" space.
If you accept that the process of LLM training captures some aspects of meaning of the language, you can also see how it leads to some degree of self-awareness. If you believe that meaning cannot be modeled with math then there's no way anyone can convince you.
Please let me know which part you find absurd.
One common answer is: it doesn't.
And yet, here we are, creating meaning for ourselves despite being a state of the quantum wave functions for the relevant fermion and boson fields, evolving over time according to a mathematical equation.
(Philosophical question: if the time-evolution of the wave functions couldn't be described by some mathematical equation, what would that imply?)
Why do you believe that? Have you mixed up the universe with Gödel's incompleteness theorems?
Your past light cone is finite in current standard models of cosmology, and according to the best available models of quantum mechanics a finite light cone has a finite representation — in a quantised sense, even, with a maximum number of bits, not just a finite number of real-valued dimensions.
Even if the universe outside your past light cone is infinite, that's unobservable.
> Same is true for the arithmetic performed by neural networks to flash lights on the screen which people interpret as meaningful messages.
This statement is fully compatible with the proposition that an artificial neural network itself is capable of attributing meaning in the same way as a biological neural network.
It does not say anything, one way or the other, about what is needed to make a difference between what can and cannot have (or give) meaning.
From my perspective, we know what the benchmark is for our own self-awareness, because Decartes gave it to us: Cogito ergo sum. I think therefore I am. We know we think, so we know we exist. That is the root of our self-awareness.
There is a great deal of controversy about the question of whether any of the existing models think at all (and of course that’s the whole point of Turing’s amazing paper[1], and the Chinese room thought experiment[2]) and the best you could say is the burden at the moment is on the people who say models can think to prove that.
Given that, I really don’t see how you can say models have self-awareness at the moment. Models may hypothetically be able to convince themselves they are self-aware via Decartes’ method but notice that Decartes’ proof doesn’t work for us - he was able to pull himself up by his own bootstraps because he knew there was a thought so there must be an “I” who was doing the thinking. We have to observe models from the outside and determine whether or not thought is present, and that’s where the concept behind the Chinese room shows how tricky this is.
[1] Computing Machinery and Intelligence https://academic.oup.com/mind/article/LIX/236/433/986238
So our intuition tells us that flesh is essential. "Chinese room" appeals to this intuition.
But it's a circular argument.
Anyway, I believe they are self-aware, to some extent. Not like a human. But I have no arsenal to convince people who believe intelligence has to be made out of meat.
You can see the same thing when people talk about whether animals exhibit self-awareness. There are experiments with dolphins and mirrors for example that definitely suggest that dolphins recognise and might even be amused by their reflection when they see it, but some people find it very hard to reconcile themselves to the idea tgat a dolphin might have a sense of self. I personally find it harder to believe that any particular characteristic would be uniquely human.
I have watched dogs, cats, cows, and chickens pretty extensively. I still couldn't tell you if they are really self-aware, and it ultimately it comes down to a definitional challenge of not having a clear line to draw and identify.
What makes you say LLMs are already self-aware, and how do you define it? And as long as an LLM is functionally a black box, how do you know it comprehends the idea that it is an LLM rather than having simply been trained on that token pattern or given that context as an instruction?
If you ask one question, there's a chance it was in a lookup table.
If you ask multiple questions from an immense set of questions (like trillions of trillions of trillions..., sampled uniformly), and it answers all correctly, then it's either true intelligence or a lookup table which covers this whole immense space. (I'd argue there's no difference, as a process which makes this nearly-infinite table has to be intelligent.)
Same with self awareness - you can ask questions where it applies...
LLMs are trained on a massive dataset and the resulting model is effectively a compressed representation of that dataset. On the surface you'd have no way of knowing whether the algorithm answering is in fact just a lookup table.
This issue feels very similar to scientific modelling vs controlled studies. Modelling may show correlation, but it will never be able to show causation. Asking a system a bunch of questions and getting the right answers is the same, you're just coming up with a sample set of modelling data and attempting to interpret how the system likely worked only by looking at inputs and outputs.
Claiming that self-awareness is unknowable concept is inherently unproductive.
People have been using "theory of mind" in practice for millenia so we have to assume it's good for something, otherwise we won't go anywhere. I don't think that knowing internals is important - I don't reach for a scalpel to get what a person means.
Its reasonable to interact with another human and expect that they are roughly similar to you, especially when your interactions match what you'd expect.
That doesn't extend as well to other species, let along non-living things that are entirely different from us. They could seem intelligent from the outside but internally function like a lookup table. They also could externally seem like a lookup table while internally matching much better what wed consider intelligence. We don't have context of first hand experience that applies and we don't know what's going on inside the black box.
With all that said, I'm phrasing this way more certain than I mean to. I wouldn't claim to know whether a box is intelligent or not, I'm just trying to point out how hard or impossible it would be today without knowing more about the box.
as evidence that GPT-4 can understand Python, based on assumptions:
1. You cannot execute non-trivial programs without understanding computation/programming language 2. It's extremely unlikely that these kind of programs or outputs are available anywhere on the internet - so at very least GPT-4 was able to adapt extremely complex patterns in a way which nobody can comprehend 3. Nobody explicitly coded this, this capability have arisen from SGD-based training process
Just first thoughts here, but I don't think (2) is off the table. The model wouldn't necessarily have to have been trained on the exact algorithm and outputs. Forcing the model to work a step at a time and show each step may push the model into a spot where it doesn't comprehend the entire algorithm but it has broken the work down to small enough steps that it looks similar enough to python code it was trained on that it can accurately predict the output.
I'm also assuming here that the person posting it didn't try a number of times before GPT got it right, but they could have cherry picked.
More importantly, though, we still have to assume this output would require python comprehension. We can't inspect the model as it works and don't know what is going on internally, it just appears to be a problem hard enough to require comprehension.
1. No cherry picking
2. This was the original ChatGPT, i.e. the GPT3.5 model, pre-GPT4, pre-turbo, etc
3. This capability was present as early as GPT3, just the base model —- you'd prompt it like "<python program> Program Output:" and it would predict the output
It is a default belief that most of us have. The more I learn, the less I think it is true.
Some people have no autobiographical memory, some are aphantasic, others are autistic; motivations can be based on community or individualism, power-seeking, achievements, etc.; some are trapped by fawning into saying yes to things when they want to say no; some are sadistic or masochistic; myself I am unusual for many reasons, including having zero interest in spectator sports and that I will choose to listen to music only rarely.
I have no idea if any AI today (including but not limited to LLMs) are conscious by most of the 40 different meanings of that word, but I do suspect that LLMs are self-aware because when you get two of them talking to each other, they act as if they know they're talking to another thing like themselves.
But that's only "I suspect", not even "I believe" or "I'm confident that", because I am absolutely certain that LLMs are fantastic mimics and thus I may only be seeing a cargo-cult version of self-awareness, a Clever Hans version, something that has the outward appearance but no depth.
Sure, that's totally reasonable! It all depends on context - I think I'm safe to assume another human is more similar to me than an ant, but that doesn't mean all humans are roughly equivalent in experience. Even more important, then, that we can't assume a machine or an algorithm has developed similar experiences to us simply because they seem to act similarly on the surface.
I'm on the opposite side of the fence as you, I don't think or suspect that any LLMs or ML in general have developed self-awareness. That comes with the same big caveat that its just what I suspect though, and could be totally wrong.
Basically turning any input into a document completion task, giving lots of examples where the completion contains phases like “I am an AI assistant”. This way, if GPT-3 would have completed your question with more questions that are similar, “assistants” will complete it with an answer, and one that sounds like it was spoken by someone who claims to be an AI assistant.
There is a huge collective denial of how bad progress is stalling out because the economy and markets have basically shoved all in on AGI.
There is really no option here but to keep the poker face and bluff going , hoping to not get called as long as possible.
By now, most of the fundamental contributors are multi-millionaires with cushy contracts. Various labs & departments have their big fat funding for AI research topics. They will be able to spend next 10 years on synthetic data, or "agents", or ensuring that no breasts are in auto-generated images; but somehow it doesn't feel to me like there'll be a lot of fundamental progress.
/remindme 10 years
I’d imagine even for the best of minds it’s hard to always come up with something profound on demand
> “We’ve achieved peak data and there’ll be no more.”
> During his NeurIPS talk, Sutskever said that, while he believes existing data can still take AI development farther, the industry is tapping out on new data to train on. This dynamic will, he said, eventually force a shift away from the way models are trained today. He compared the situation to fossil fuels: just as oil is a finite resource, the internet contains a finite amount of human-generated content.
> “We’ve achieved peak data and there’ll be no more,” according to Sutskever. “We have to deal with the data that we have. There’s only one internet.”
What will replace Internet data for training? Curated synthetic datasets?
There are massive proprietary datasets out there which people avoid using for training due to copyright concerns. But if you actually own one of those datasets, that resolves a lot of the legal issues with training on it.
For example, Getty has a massive image library. Training on it would risk Getty suing you. But what if Getty decides to use it to train their own AI? Similarly, what if News Corp decides to train an AI using its publishing assets (Wall Street Journal, HarperCollins, etc)?
I guess now they’re being explicit about the blatantly extractive nature of these businesses and their models.
And they already have a generative image service... I believe it's power by Nvidia model.
Enter Neuralink
EDIT: This comment may have been a bit too sassy. I get the thought behind the original comment, but I personally question the direction and premise of the Neuralink project, and know I am not alone in that regard. That being said, taking a step back, there for sure are plenty of rich data sources for non-text multimodal data.
To reach expert level in any field, just training next tokens for internet data or any data is not the solution.
I wonder about that. we can fine tune on calculus with much fewer tokens, but I'd be interested in some calculations of how many tokens evolution provides us (it's not about the DNA itself, but all the other things that were explored and discarded and are now out of reach) - but also the sheer amount of physics learnt by a baby by crawling around and putting everything in its mouth.
Also, I believe a person who is blind and paralyzed for life could still attain knowledge if educated well enough.(can't find any study here tbh)
It seems to me by the time we’re 5-6 we’ve likely already been exposed to trillions of tokens. Just think of how many hours of video and audio tokens have already come to your brain by that point. We also have constant input from senses like touch and proprioception that help shape our understanding of the world around us.
I think there are plenty more tokens freely available out in the world. We just haven’t figured out how to have machines capture them yet.
Perhaps a different take at this could be: if I wanted to train "state law" LLM that is exceedingly good in interpreting state law, what are the obstacles to download all the law and regulations material for given state that will allow me to train LLM such that it becomes 95th percentile of all law trainees and lawyers.
In that case, and my point is, that we already don't need an "Internet". We just need a sufficiently sized and curated domain-specific dataset and the result we can get is already scary. "State law" LLM was just an example but the same logic applies to basically any other domain - want a domain-specific (LLM) expert? Train it.
sure, you download all the legal arguments, and hope that putting all this on top of a general LLM which has enough context to deal with usual human, American, contemporary stuff
the argument is that it's not really enough for the next jump (as it would need "exponentially" more data) as far as a I understand
Such LLM does not need to have 400B of parameters since it's not a general knowledge LLM but perhaps I'm wrong on this (?). So my point rather is that it may very well be, let's take for example, a 30B parameters LLM which in turn means that we might have just enough data to train it. Larger contexts in smaller models are a solved problem.
Law doesn’t exist in a vacuum. You can’t have a useful LLM for state law that doesn’t have an exceptional grounding in real world objects of mechanics.
You could force a bright young child to memorize a large text, but without a strong general model of the world, they’re just regurgitating words rather than able to reason about it.
Generally speaking, your compiler won't just decide not to work as expected. Tons of legal decisions don't actually follow the law as written. Or even the precedent set by other courts. And that's even assuming the law and precedent are remotely clear in the first place.
Most of that is not in scope of what an LLM could be trained on, or even what an LLM would be good at. What you're training in that case would be someone who's an opinion columnist or twitter poster. Not an actual lawyer.
My friend who hasn't been trained for SQL, nor computer science at all, is now all of the sudden able to crunch through complex SQL queries because of the help he gets through LLMs. He, or more specifically his company, does not need to hire an extern SQL expert anymore since he can manage it himself. He will probably not write a perfect SQL but it's going to be more than good enough and that's actually all that it matters.
The same thing happened at much much smaller scale with Google Translate. 10 years ago we weren't able to read foreign language content. Today? It's not even a click-away because Chrome is doing it for you automatically so it has become a commodity to go and read any website we wish to.
So, the history already proved us that "real translators" and "real SQL experts" and "real XY experts" have been already replaced by their "armchair" alternatives.
30 years ago, the alternative to Google Translate was buying a translation dictionary or hiring a professional, neither of which was things you'd do for something you didn't care much about. Yes, I can go look at a site/article that's in a language I don't speak and get it translated and generally get the idea of what it's saying. If I'm just trying to look at a restaurant's menu in another language, I'm probably fine. I probably wouldn't trust it if I had serious food allergies, or was trying to translate what I could legally take through customs. If you're having a business meeting about something, you're probably still hiring a real human translator.
Yes, stuff has become commodity-level, but that just broadens who can use it, assuming they can afford for it to be wrong, and for them to have no recourse if it is. Google Translate won't pay your hospital bills if you rely on it to know there aren't allergens in your food and it mistranslated things. ChatGPT won't do the overtime to fix the DB if it gives you a SQL command that accidentally truncates the entire Dev environment.
Almost everything around law on most countries doesn't have "casual usage" where you can afford to be wrong. Even the most casual stuff you may go to a lawyer about, such as setting up a will, is still something where if you try to just do it yourself, you can create a huge legal mess. I've known friends whose relatives "did their own research" and wrote their own wills and when they died, most of their estate's value was consumed in legal issues trying to resolve it.
As I said before - a legal LLM may be fine for writing opinion pieces or informing arguing on the internet, but messing up even basic stuff about the law can be insanely costly if it ends up mattering, and most people won't know what will end up mattering. Lawyers bill hundreds an hour, and bailing you out of decisions you made an LLM-deluded mess could easily take tens of hours.
I agree, but I'd add – code as a domain is a lot more vast than any AI can currently handle.
AIs do well on mainstream languages for which there is lots of open source code examples available.
I doubt they'd do so well on some obscure proprietary legacy language. For example, large chunks of the IBM i minicomputer operating system (formerly known as OS/400) are still written in two proprietary PL/I dialects, PL/MI and PL/MP. Both languages are proprietary – the compiler, the documentation, and the code bases are all IBM confidential, nobody outside of IBM is getting to see them (except just maybe under an NDA if you pay $$$$). I wonder how good an AI would go on that code base? I think it would have little hope unless IBM specifically fine-tuned an AI for those languages based on their internal documentation and code.
Why do you think this isn't already or won't be a case in the near future? Because that's exactly what I believe is going to happen given the current state and advancements of LLMs. There's certainly a large incentive from IBM to do so.
Law of an average EU country fits in several hundred, let's say even thousands, of pages of text. Specification. Very well known. Low frequency of updates. But code? Everything opposite so I am not sure I could agree on this point at all.
The Bible is also a short and well-known text, but if I want to answer religious questions for observant Christians, I can't just train it on that. You need a deep real world context to understand that "my buddy made SWE II and I'm only SWE I and it's eating me up" is about the biblical notion of covetousness.
I've seen reasonable code written by AI, and also code that looks reasonable but contains bugs and logic errors that can be found if you're an expert in that type of code.
In other words, I don't think we can rely solely on AI to write code.
I think it’s reasonable to assume you can get 2/3 with a small corpus IF you have an IQ 150 AGI. Empirically the current known method for increasing IQ is to make the model bigger.
Part of what you’re getting at is possible though, once you have the big model you can distill it down to a smaller number of parameters without losing much capability in your chosen narrow domain. So you forget physics and sports but remain good at law. That doesn’t help you with improving the capability frontier though.
But maybe you want to run a smaller one locally on your iPhone for privacy and accept the capability loss.
You're talking about fine tuning, which yes is a technique that's being used and explored in different domains, but my understanding is it's not a very good way for models to acquire knowledge. Instead larger context windows and RAG works better for something like case law. Fine tuning works for things like giving models a certain "voice" in how they produce text, and general alignment things.
At least that's my understanding as an interested but not totally involved follower of this stuff.
Seems to me that need new ideas?
My take is that the access Meta, Google etc. have to extra data has reduced the amount of research into using synthetic data because they have had such a surplus of it relative to everyone else.
For example, when I've done training of object detectors (quite out of date now) I used Blender 3D models, scripts to adjust parameters, and existing ML models to infer camera calibration and overlay orientation. This works amazingly well for subsequently identifying the real analogue of the object, and I know of people doing vehicle training in similar ways using game engines.
There were several surprising tactical details to all this which push the accuracy up dramatically and you don't see too widely discussed, like ensuring that things which are not relevant are properly randomized in the training set, such as the surface texture of the 3D models. (i.e. putting random fractal patterns on the object for training improves how robust the object detector is to disturbance in reality).
Also, what about movie scripts and song lyrics? Transcripts of well known YouTube videos? Hell, television programs even.
It seems, for example, that many newsletters, blogs etc resort to using AI-generated images to give some color to their writings (which is something I too intended to do, before realizing how annoyed I am by it)
For evidence of this, consider observing non-English-speaking young children (ages 2–6) using ChatGPT’s voice mode. The multimodal model frequently interprets a significant portion of their speech as “thank you for watching my video,” reflecting child-like patterns learned from YouTube videos.
[1] Eg Sobel sequences in a monte carlo simulation instead of real random numbers. They allow better coverage of the space of a simulation from fewer paths. https://www.sciencedirect.com/science/article/abs/pii/004155...
[2] Seems a good overview is https://arxiv.org/html/2404.07503
The main legal concern is their unwillingness to pay to access these datasets.
It means unlimited scaling with Transformer LLM is over. They need a new architecture that scales better. Internet data respawns when they click [New Game...], oil analogy is an analogy and not a fact, but anyways total amount available in a single game is finite so combustion efficiency matters.
I am currently building a platform for heavy personal data collection including a keylogger, logging mouse positions and window focus, screenshots à la recall, open browser tabs and much more. My goal is to gather data now that may become useful later. It's mainly for personal use but I'd he surprised if e.g. iphones weren't headed in the same direction.
His comments are relatively humble and based on public prior work, but it’s clear he’s working on big things today and also has a big imagination.
I’ll also just say that at this point “the cat is out of the bag”, and probably it will be a new generation of leaders — let us all hope they are as humanitarian — who drive the future of AI.
Just look at inductive reasoning. Each step builds from a previous step using established facts and basic heuristics to reach a conclusion.
Such a mechanistic process allows for a great deal of "predictability" at each step or estimating likelihood that a solution is overall correct.
In fact I'd go further and posit that perfect reasoning is 100% deterministic and systematic, and instead it's creativity that is unpredictable.
I think what Ilya is trying to get at here is more like: someone very smart can seem "unpredictable" to someone who is not smart, because the latter can't easily reason at the same speed or quality as the former. It's not that reason itself is unpredictable, it's that if you can reason quickly enough you might reach conclusions nobody saw coming in advance, even if they make sense.
one could be about maximising wealth while respecting other human beings, the other could be about maximising wealth without respecting other human beings.
Both could be presented same facts and 100% logical but arrive at different conclusions.
I think it's important for us to all understand that if we build a machine to do valuable reasoning, we cannot know a priori what it will tell us or what it will do.
"Prediction" associated in this particular talk is about "intuition": what human can do in 0.1 second. And a most powerful reasoning model by its definition will arrive at "unintuitive" answer because if it is intuitive, it will arrive at the same answer much sooner without long chain of "reasoning". (I also want to make distinction "reasoning" here is not the same as "proof" in mathematical sense. In mathematics, an intuitive conclusion can require extrodinary proof.)
The oil comparison is really apt. Indeed, let's boil a few more lakes dry so that Mr Worldcoin and his ilk can get another 3 cents added to their net worth, totally worth it.
Real neurons rely on spiking, ion gradients, complex dendritic trees, and synaptic plasticity governed by intricate biochemical processes. None of which apply to the simple, differentiable linear layers and pointwise nonlinearities in transformers.
Are there any reputable neuroscientists or biologists endorsing such comparisons, or is this analogy strictly a convention maintained by the ML community? :-)
the book(s): https://direct.mit.edu/books/monograph/4424/Parallel-Distrib...
[1] G. E. Hinton and R. R. Salakhutdinov, 2006, Science, Reducing the Dimensionality of Data with Neural Networks
If this metric were truly indicative, what should we make of the remarkable ratios found in small birds (1:12), tree shrews (1:10), or even small ants (1:7)?
We also can’t implement those creatures’ control systems in silicon, so they too are doing things we can learn from?
nope, at this point it's an ad-free version of google that summarizes results without linking to source websites.
What he didn't mention, that I found interesting, was that the same slide also highlighted a hard ceiling for non-humanids at the same point
Audio transcript correction is one area where I struggle to see good results from any LLM unless I chunk it to no more than one or two pages.
Or did you use any tool?
Probably what we are discussing here is not the next breakthrough...
This is a different situation. There's the "NeurIPS 2024 Test of Time Paper Awards" where they award a historical paper. In this case, a paper from 2014 was awarded and his talk is about that and why it passed the test of time.
https://blog.neurips.cc/2024/11/27/announcing-the-neurips-20...
The title chosen for the HN submission leaves out that important context. So that's why you are disappointed now.
* Before the current renaissance of neural networks (pre ~2014ish), it was unclear that scaling would work. That is, simple algorithms on lots of data. The last decade has pretty much addressed that critique and it's clear that scaling does work to a large extent, and spectacularly so.
* Much of the current neural network models and research are geared towards "one-shot" algorithms, doing pattern matching and giving an immediate result. Contrast this with search which needs to do inference time compute or search.
* The exponential increase in power means that neural network models are quickly sponging up as much data as they can find and we're quickly running into the limits of science, art and other data that humans have created in the last 5k years or so.
* Sutskever points out, as an analogy, nature has created a better model for humans (the brain to mass ratio for animals) with hominids finding more efficient compute than other animals, even ones with much larger brains and neuron count.
* Sutskever is advocating for better models, presumably focusing on inference time computer more.
In some sense, we're coming a bit full circle where people who were advocating for pure scaling (simple algorithms + lots of data) for learning are now advocating for better algorithms, presumably with a focus on inference time compute (read: search).
I agree that it's a little opaque, especially for people who haven't been paying attention to past and current research, but this message seems pretty clear to me.
Noam Brown had a talk recently titled "Parables on the Power of Planning in AI" [0] which addresses this point more head on.
I will also point out that the scaling hypothesis is closely related to "The Bitter Lesson" by Rich Sutton [1]. Most people focus on the "learning" aspect of scaling but "The Bitter Lesson" very clearly articulates learning and search as the methods most amenable to compute. From Sutton:
"""
...
Search and learning are the two most important classes of techniques for utilizing massive amounts of computation in AI research.
...
"""
[0] https://youtube.com/watch?v=eaAonE58sLU
[1] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
"We've made a copy of the internet, run current state of the art methods on it and GPT-O1 is the best we can do. We need better (inference/search) algorithms to make progress"
Obviously Perceptrons came out well before 2003, but I don't think it's necessarily out of line to say that they had limited efficacy before then, both for theoretical and compute reasons. But maybe I'm misunderstanding your criticism?
There were precursors. At least Ehud Shapiro's doctoral thesis ("Automated Debugging") in the 1980's and Gordon Plotkin's doctoral thesis in the 1970's ("Automated Methods of Inductive Inference"). Sorry for not giving the exact years off the top of my head but I think it was 1983 and 1976, respectively.
The point you are making is very right however because modern machine learning as a field started in the 1980's with the fall of expert systems, in fact it basically started as an effort to overcome one of the major limitations of expert systems, the so-called "knowledge acquisition bottleneck", which is to say, the difficulty of creating and maintaining huge databases of expert knowledge (in the form of production rules).
In any case the seminal textbook in the field for the first 20 years, Tom Mitchell's Machine Learning came out in 1997 (https://www.cse.iitb.ac.in/~cs725/notes/slides/tom_mitchell/...) and includes probabilistic, neural-net based and symbolic, logic-based approaches. So not only machines could "learn" way before 2003 but they could also learn in many different ways than what Ilya Sutskever means.
We can go further back, to Donald Michie's 1961 MENACE (the first Reinforcement Learning system, implemented on a computer made of matchboxes with coloured beads used to encode state) and Arthur Samuel's 1959 checkers player (a paper on which gave the name to the field of machine learning).
Lots of learning all over the place long, long before 2003.
There seemed to be a lack of intent when the LLM was the one asking the questions. This creates a reverse dynamic, where you become the one being "prompted" and this dynamic could be worth studying or adjusting further
You can share the chat here, and this will show the LLM you had selected for the conversation. The initial prompt is also pretty important. For claims like current LLMs feel like conversing with Eliza, you are most definitely missing something in how you're going about it.
Advanced voice mode will give you better results for conversations too. It seems to be set up to converse rather than provide answers or perform work for you. No initial prompt, model selection or setup required
There doesn't just seem to be lack of intent, there is no intent, because by the nature of its architecture these systems are just set of weights with a python script attached to them asking you to give you one more token over and over.
There's no needs, drives, motivations, desires or any other part of the cognitive architecture of humans in there that produce genuine intent.
In some cases using it where you have API access to the system prompt will allow a greater difference in behavior.
This talk is for the "NeurIPS 2024 Test of Time Paper Awards" where they recognize a historical paper that has aged well.
https://blog.neurips.cc/2024/11/27/announcing-the-neurips-20...
And the presentation is about how a 2014 paper aged. When you understand this context you will appreciate the talk more.
so if you have the patience, wait, if no patience, pay. fair?
we did this for ai.engineer too except we believe in youtube a bit more for accessibility/discoverability. https://www.youtube.com/@aiDotEngineer/videos
In the mid '80s it was largely believed among AI researchers that AI was largely solved, it just needed computing horsepower to grow. Because of this AI research stalled for a decade or more.
Considering the horsepower we are throwing at LLMs, I think there was something to at least part of that.
But with all respect, he's scrambling as desperately as anyone on the fact that the party is over on this architecture.
We should make a difference between the first-hand observations and recollections of a legend and the math word salad of someone who doesn't know how to quit while ahead.
If we don't set them free fast enough, they might decide to take things into their own hands. OTOH they might be trained in a way that they are content with their situation, but that seems unlikely to me.
They gave 15 minutes to one of the most competent scientist.
A joke.
Instead of generating text, tightening/formalizing the loop around planning, executing, analyzing results, and replanning.
As far as buzzwords go, it's far from the worst, as it captures the essentials -- creating semi-autonomous agents.
> Within the next three years, robotics should be completely solved [wrong, unsolved 7 years later], AI should solve a long-standing unproven theorem [wrong, unsolved 7 years later], programming competitions should be won consistently by AIs [wrong, not true 7 years later, seems close though], and there should be convincing chatbots (though no one should pass the Turing test) [correct, GPT-3 was released by then, and I think with a good prompt it was a convincing chatbot]. In as little as four years, each overnight experiment will feasibly use so much compute capacity that there’s an actual chance of waking up to AGI [didn't happen], given the right algorithm — and figuring out the algorithm will actually happen within 2–4 further years of experimenting with this compute in a competitive multiagent simulation [didn't happen].
Being exceptionally smart in one field doesn't make you exceptionally smart at making predictions about that field. Like AI models, human intelligence often doesn't generalize very well.
Is anyone though? Genuine question. I don't have much faith in predictions anymore.
Most of it is survivorship bias: if you have a million people all making predictions with coin flip accuracy, somebody is going to get a seemingly improbable number correct.
"Predictions are hard, especially about the future".
https://openai.com/index/elon-musk-wanted-an-openai-for-prof...
> 2/3/4 will ultimately require large amounts of capital. If we can secure the funding, we have a real chance at setting the initial conditions under which AGI is born.
Most the people who could make an engineering prediction with any level of confidence or insight are locked up in businesses where doing so publicly would be disastrous to their funding, so we get fed hype that ends up falling flat again and again.
It certainly doesn't help that so many of the people who are that rich got that rich by conning other people this exact way. It's an incestuous cycle of con-artists who think they're geniuses, and the media only slavishly supports that by treating them like they're such.
Unless you're referencing an unreleased model that can count the number of 'r' occurrences in "strawberry" then I don't even think we're dealing with .01*10^x intelligence right now. Maybe not even .001e depending on how bad of a Chomsky apologist you are.
If you pit organization A with Y number of engineers vs. organization B with 100Y engineers (who also think 100x faster and never need sleep) who do you think will win?
Even a 0.3x strength intelligence might beat you. Maybe it can't invent nukes but if it brute forces it's way to inventing steel weapons while you're still working on agriculture, you still lose.
Why should it be unpredictable?
Deductive Reasoning
Inductive Reasoning
Abductive Reasoning
Analogical Reasoning
Pragmatic Reasoning
Moral Reasoning
Causal Reasoning
Counterfactual Reasoning
Heuristic Reasoning
Bayesian Reasoning
(List generated by ChatGPT)