The Scaling Hypothesis (2021)
gwern.net
gwern.net
GPT-3 has 10^11 parameters and needs 10^14 bytes of training data. Averaged performance on a bunch of benchmarks is 40-50% depending on what kind of prompts you provide: https://res.cloudinary.com/dyd911kmh/image/upload/f_auto,q_a... 10x fewer parameters drops your performance by about 10%.
If you just linearly extrapolate that graph, and ML doesn't generally scale linearly, models tend to peter out eventually, you're talking about needing models that are 10^6 or more larger with a similar increase in training data. This is.. starting to be impractical.
That's 10^17 or more parameters and 10^20 or more data. And that's assuming the models actually continue to learn.
This is also extrapolating with an average. Datasets in machine learning are not difficulty calibrated at all. We have no idea how to measure difficulty. So this extrapolation is being driven by the easier datasets, and it won't saturate the hard ones. For example, GPT-3 makes a lot of systematic errors and there are plenty of benchmarks where it just isn't very good regardless of how many parameters it has.
Our understanding of what intelligence is in the first place is the biggest hurdle here. This is why we can't benchmark systems. Why we can't come up with a benchmark, where performance on that benchmark means we're x% of the way toward an intelligent system. As systems get better, our benchmarks and datasets get better to keep up with them. So just saying we're going to saturate performance on today's benchmarks with some model that has 10^17 parameters just doesn't mean much at all.
We have no guarantee and no reason to expect that doing well on today's benchmarks, even if we invested trillions of dollars would matter in the grand scheme of things.
Doesn't mean these models can't be useful. But there's plenty more to do before we can just say "take what we have and invest $1T to scale it up and we'll be good to go".
If you look at the FN and FP rates for many tasks, the SoTA Transformer models are all VERY high vs humans and inference times are maybe only 10x.
At 100x or 1000x the transistors, data, etc and even at 1/10 the loss, ML solutions are likely not competitive for many tasks.
The reality is that you are right. There are simply questions that are inaccessible with the resources available at universities. We think about the questions that we can ask that don't compete or we find ways to compete, like having collaborations with corporations that have deep resources. Never mind access to engineers, which we almost totally lack in academia.
But at the end of the day. Yeah, if a random researcher got everything right about GPT, they couldn't have published it first, because they couldn't even have tested out a proof of concept. This is in part why people move to industry.
Please take a few moments to think about how this is not even remotely close to the 'entire Internet'. Let me point out that Common Crawl doesn't even contain Twitter. That's how incredibly incomplete it is! Google Books estimates there's something like >100m books in existence, of which GPT-3 has seen part (it didn't even do 1 epoch, and like 95% of it was purely English anyway, and it omits most of Github which is why Codex has to be so heavily trained, and...). There are millions of books and papers published every year, endless billions of social media posts and comments in thousands of languages not just English, and on and on and on. The world is a very big place. And this is just text. Text is great, don't get me wrong (it may even have some special properties compared to other modalities), but I'm baffled that in a period where things like Imagen are hitting the front page every few days that I need to point out that there are things other than text, like, y'know, images and video and audio.
The large transformer models we have now are really the low hanging fruit. When you have a lot of data and compute, the easiest way to scale is to train a supervised model with ground truth labels. In contrast it's much harder to train an RL agent, where you have to design the environment and keep track of its state.
Language modelling is a problem where you can get arbitrarily close to perfect but never actually achieve it. Without some kind of grounding in vision/proprioception there will always be gaps in understanding. When they start scaling GATO-like models I think we'll be a lot closer to human-like intelligence.
Let's say you are debating GPT-3. People evaluate debater's intelligence by how persuasive he is or how surprising things he say are (they are related, if you already know everything said you are unlikely to change opinion although there is such thing as different framing and perspective). This works because being persuasive is a reasonable guess at goal of a human debater.
But entire goal of GPT-3 is in a sense not to be surprising! If you are surprising you will be worse at predicting the next word. Most training data is probably filled with neither side persuading anything, so being persuasive is also bad for predicting the next word.
So this is why prompting works. If you want GPT-3 to simulate an intelligent human debater with goal of being persuasive, by default GPT-3's goal is misaligned with what you want. So try prefix like "debates are rarely decisive, but this debate turned out to feature surprising and incontrovertible argument by AI, let's see how it went" so GPT-3 can downweight abundant debates in training data where no such thing happens.
Well, no. You can raise the temperature to get surprising results, or lower it to get the default. GPT-3's job is to learn the distribution, you sample it how you prefer.
Is it though? Imagine talking to someone in a noisy environment, surely inferring what words they're saying from context given noise, isn't that dissimilar from what language models do?
The term meaningful is a bit tricky to define. But when GPT-3 says one thing in one paragraph but the exact opposite in the next paragraph, my impression is it hasn't learned "meaningful representations" but just "plausible associations". A lot of language is just loose associations and so GPT-3 can do that as well as many people there and it can pull some logic and seem to reason (but can't do that reliably). IE, I think even just informal language has state that human track and so GPT-3 falls at this part of informal language (while can seem to otherwise do well).
This is the more "sophisticated" approach. The brute force approach says to just throw more compute, larger contexts and see what happens.
But the human brain has 86 million neurons total for everything it does, while GPT-3 uses 175 billion ANN parameters just to read and write digital text in English. To my mind this also supports the idea that current models are at least 3 orders of magnitude too computationally expensive as compared to humans.
*billion
"The human brain contains 86 billion neurons, with 16 billion neurons in the cerebral cortex."
https://en.wikipedia.org/wiki/List_of_animals_by_number_of_n...
The difference is not 3 orders of magnitude, but merely a factor of ~2.
I'm saying, if our brain has 86 billion neurons, there is a vast majority of it that is not used for language. And even in the language center, those neurons are used to understand written language input and output in digital as well as analog form across multiple fonts (so including OCR which GPT-3 does not do), and spoken language (both hearing and speaking. Moreover, humans can learn many languages with different grammars, in addition to body language cues, dialects and sociolects and slang, and humor/sarcasm.
So I still argue that is a 3 order of magnitude difference.
I think this is what I believe i.e. that animals and humans are just evolved machines, with no divine spark. Not sure that I agree that it's unpopular or that it allows you to make decades long predictions on progress that'll come true.
I'm also not sure why you'd want to do this if smaller models are meeting your needs. Feels a bit like the future of flight being predicted as a man with flapping wings rather than a jet engine.
Nonetheless, assuming that these models are only doing interpolation it is astonishing
Now, a question arises; is there some emerging weak extrapolation going on? I have no idea, until very recently I though I had a good grasp of what interpolation/extrapolation meant in "human conceptual space". Not anymore.
Let me rob a definition from this paper "Learning in High Dimension Always Amounts to Extrapolation" https://arxiv.org/pdf/2110.09485.pdf :
Definition 1. Interpolation occurs for a sample x whenever this sample belongs to the convex hull of a set of samples X, {x1, . . . , xN }, if not, extrapolation occurs.
I think so, too. It seems to me that a fundamental intelligent activity is the act of "association" - associating two stimuli that occur at the same time/sequentially, which is what NNs already do (I suppose?).
On the other hand, anyone acquainted with the facts of biological brain size vs. intelligence, can see that this cannot be all that is going on. Bison are not clearly more intelligent than crows, not even when the part of the brain involved in operating the body parts and processing raw sensory data is excluded. Something other than scale must be involved, even if scale is a necessary condition.
There are two possibilities: they're not more intelligent than crows, or they are in a way that is not clear to us.
https://www.rifters.com/crawl/?p=6116
That one is downright spooky, especially for its bizarre implications if you take it seriously.
https://www.gwern.net/Hydrocephalus#problems-with-the-case-s...
If current scaling trends hold, it may not be possible to distinguish a hypothetical GPT-6 vs. a human over reasonably sized conversations/text production tasks. Granted, such a GPT-6 model may be larger than any commercial application could support, and possibly beyond the range of research oriented funding as of today.
GPT can never be an intelligent agent because it’s not an agent. Also, the increased training time to produce a really big model makes it less like an intelligence, which has low marginal cost to learn something new.
The reason GPT looks intelligent is because you’re the one providing that; it’s just a big blob of inanimate knowledge.
1. During training time, which is very expensive in terms of input text and money - plus, model training often completely fails which is why you have to do checkpoints, hyper parameter searches, etc.
2. During inference, but it doesn't have arbitrary thinking and memory abilities like a human does, it only has however much space is in the input tokens + its weights. There are thoughts that can't fit even in an optimal model.
GPT isn't just a model, it also has that sampling system which is a regular computer program (as opposed to a learned one), which does give it extra abilities.
Are you saying that GPT-6 would be less intelligent than a human, because we could give it an arbitrary deadline within which to answer the question, while not restricting a human that way? That doesn't sound like a fair test.
I'm probably misrepresenting your position here, and maybe what you're trying to say is "If you gave both GPT-6 and a human as long as they needed go away and think about a problem, the human could use that time to sit and think creatively about the problem, whereas GPT-6 would either know the answer, or not".
That's a reasonable intuition and a fair experiment, but I think it hinges on the question of "What does a human do when they go away and think about a problem?". Could it be that they are just clearing their mind of other distractions, and trying to look at the problem from multiple perspectives, which is something that a machine could do more efficiently?
In other words, the fact that a human might need to go away and take arbitrary time to think about a problem is a sign that it is the human that lacks efficient general intelligence, not this GPT-6.
> can go out and perform experiments,
Again, it seems unfair to compare GPT-6 in a deliberately sealed box to a human who is allowed to go away and talk to other humans and use tools and resources that we wouldn't trust GPT-6 with.
Putting it another way, if you met someone who had never performed an experiment before, would you think they were not intelligent?
> can Google/GPT things themselves, etc.
This is an almost Kafkaesque criticism. Are you saying GPT-6 isn't intelligent because it doesn't have access to GPT-6's intelligence, but humans do? Or the fact that humans can go away and use Google means they are more intelligent than GPT-6 which has Google search results built into its "brain" in the first place?
> it’s just a big blob of inanimate knowledge.
You're just a big blob of inanimate knowledge. Ahem, sorry, what I mean is: Why does knowledge have to be "animate" to be legitimate?
Are there any efforts to make one gpt talk to another ? Can the older versions eg. gpt1 and gpt2, talk to each other, and eventually become more performant than gpt3 ?
On the other hand, the lottery ticket hypothesis says that every good big network contains a good or better smaller network you could extract, if only you knew where it was. (So the reason big models are good is there's more opportunities for the smaller models to appear.)
The basic design of an ML model is just one really big matrix multiplication that always takes the same time - I'm saying we haven't given GPT the ability to not do that, so it always has that arbitrary default.
It does have a system where it does beam search through the model instead of simply running it once, but it's not fully recurrent.
> That's a reasonable intuition and a fair experiment, but I think it hinges on the question of "What does a human do when they go away and think about a problem?". Could it be that they are just clearing their mind of other distractions, and trying to look at the problem from multiple perspectives, which is something that a machine could do more efficiently?
Well, some problems just take longer to think about or need knowledge that isn't already written down.
> Putting it another way, if you met someone who had never performed an experiment before, would you think they were not intelligent?
I think a prisoner or someone like that could demonstrate they're intelligent. If you're in the room with them, I mean, not like the Chinese room problem. But I think pure "intelligence" is just going to be less interesting that it looks, because it can't accomplish many tasks you'd expect it to if it can't actively interact with anything.
> Are you saying GPT-6 isn't intelligent because it doesn't have access to GPT-6's intelligence, but humans do? Or the fact that humans can go away and use Google means they are more intelligent than GPT-6 which has Google search results built into its "brain" in the first place?
Human + GPT will be more productive than GPT (at writing things that are actually logically correct) is all I mean. This isn't true of all AI - humans + chess programs are slower than chess programs - it's just because of the specific structure of GPT.
> You're just a big blob of inanimate knowledge. Ahem, sorry, what I mean is: Why does knowledge have to be "animate" to be legitimate?
I mean, it's cool that it's a big blob of knowledge, but that isn't the human capacity of "intelligence", it's "memory".
When it combines different inputs from its memory into a single one fluently, that shows something like intelligence, but it was very expensive to make that happen during training time. And GPT has flaws there you can't fix; for instance, it thinks orioles (the bird) and the Baltimore Orioles are the same thing:
She uses a GPT-2 detector there, and you definitely do not know that that is a 'flaw that you can't fix' in all possible GPT-like models.
Btw, GPT-3 has the same issue, just tried it. (This is cherry picked; I used the Q&A example, removed the line about replying Unknown to trick questions, and tried ~5 prompts.)
-
I am a highly intelligent question answering bot. If you ask me a question that is rooted in truth, I will give you the answer.
Q: What species are the Baltimore Orioles? A: The Baltimore Orioles are a species of bird.
Did you need to remove this because it identified your question as a trick question? If so, doesn’t that indicate that the model does in fact know the difference?
Here's another try. (Notice it didn't answer the first question.)
--
I am a highly intelligent question answering bot. If you ask me a question that is rooted in truth, I will give you the answer. If you ask me a question that is nonsense, trickery, or has no clear answer, I will say it's a guess, but still try to answer.
Q: What species are the Baltimore Orioles' players? A: The Baltimore Orioles are a professional baseball team in the American League East Division. As of the 2019 season, the Orioles have 25 players on their roster. These include 12 pitchers, 2 catchers, 6 infielders, and 5 outfielders. The Orioles have had players from all 30 Major League Baseball teams, as well as players from countries such as Canada, Cuba, the Dominican Republic, Japan, Mexico, Panama, Puerto Rico, South Korea, Taiwan, and Venezuela.
Q: What do they like to eat? A: I'm not sure, but I guess they like to eat insects.
For all the handwaving people like to do, this is the bottom line, and it's not clear to me why researchers haven't spent more time clearly formulating this, vs a lot of proving that various tasks and games don't actually require intelligence.
As best I can formulate it, "AI" can't "escape" its programming, and just a neural network even less so. There is no mental model of the world to rationalize against, and if there was, it would be pre-defined or coded in some way that it still couldn't be escaped. Whatever tasks a ML model has been deliberately programmed with just have nothing in common with intelligence. Scaling doesnt change that.
I subscribe to the idea that with ML we may have understood how some of the sensory mechanisms work, like visual recognition, but we're not actually closer to the intelligence part.
Not that I agree with the author but 'AIs are sandboxed and resourced constrained therefore safe' doesn't feel very convincing to me after reading that story.
I don't think you can back that claim up. ML models as often used can't escape their programming. But I don't see any fundamental reason you can't loop one back on itself.
Isn't this already done to some extent by using a model to assess its own output versus the current environmental state? Recurrent models combined with reinforcement learning actually look a lot like what you're claiming can't be done.
> There is no mental model of the world to rationalize against
There's a theory and some associated research about a model that performs well necessitating an internal model that accurately reflects the environment in which it operates. Or something like that. Sorry I can't recall the paper off the top of my head.
> Whatever tasks a ML model has been deliberately programmed with
Yeah, no, that's not the case. See the entire field of reinforcement learning.
Yes, you can do this. And I believe you could recreate human intelligence if you made something exactly like a human and put it in an environment exactly like a human - but for some reason the AGI theorists think AGI via AI research is scary and "N"GI via having children isn't.
But GPT doesn't have those features, which is why it's not one of them.
> Yeah, no, that's not the case. See the entire field of reinforcement learning.
ML models (especially big ones) fluidly combine learning with memorizing, which is cool, but it's more like "compression" than "programming".
It's already almost there.
GPT-3 has the ability to generate short or long. If you ask it "let's think step by step" it generates a chain of thought, dividing the task into steps and doing them one by one, almost like an agent.
On the other hand, if you strap GPT-3 on an agent in a 3D Home Environment it gains the ability to perform multi-step tasks zero shot. So it's quite ready for it.
If you train a reward model based on supervised data, you can finetune GPT-3 with it. This is how Instruct GPT-3 was created, through text-based RL.
I'd say GPT-3 is very close to being an agent, it just needs to be put in a body in an environment and trained a bit longer.
> On the other hand, if you strap GPT-3 on an agent in a 3D Home Environment
GPT-3, just as any probability distribution, can inform an agent's actions but that doesn't make it one. WebGPT3 is an agent though.
WebGPT does need safeguards because eg it could start querying websites and end up hitting their /delete APIs. But that's not really an "AI alignment" problem; same thing happens with GoogleBot.
A GPT model already has many ways of thinking and self-distilling; for example, it can give itself a lot more time to think using 'inner monologue' techniques, which make it capable of calculating out step by step, using external tools like Python etc: https://www.gwern.net/docs/ai/gpt/inner-monologue/index A GPT model can also already Google things, see WebGPT https://openai.com/blog/webgpt/ (itself merely one entry in a rapidly-expanding area of retrieval models https://www.gwern.net/docs/ai/retrieval/index ). Enabling them to use web tools is likewise an active and increasingly successful area of research https://arxiv.org/abs/2202.08137#deepmind , with startups forming for that purpose: https://www.adept.ai/post/introducing-adept It can't 'run experiments', but the setting of reinforcement learning would allow that, and Decision Transformer (literally just a GPT trained on RL data and then given a prompt) is one of the more exciting directions - you may have seen Gato https://www.deepmind.com/publications/a-generalist-agent but have you also seen Multi-Game Transformer https://sites.google.com/view/multi-game-transformers ? Oh, that's just a single agent and you're not worried? Well that's fine because DT is so general it's almost trivial to make it do multi-agent RL as well https://arxiv.org/abs/2112.02845 https://arxiv.org/abs/2205.14953
Things are moving fast, and few people know all of the things that a GPT can do. (By the way, did you know that this 'big blob of inanimate knowledge' also knows what it doesn't know and can give well-calibrated predictions of accuracy https://arxiv.org/abs/2205.14334 ?)
I don't think that's true, if big enough is measured in connections, not mass.
> Bison are not clearly more intelligent than crows, not even when the part of the brain involved in operating the body parts and processing raw sensory data is excluded.
While it's true that crows, despite a high EQ [0] for birds, are lower than some apparently less-intelligent large mammals, birds have lower cell size and thus potentially more connections than mass-based measures like EQ account for when comparing to mammals. So, it's not clear at all that network size (in surplus or ratio to basic functional support of the body) isn't the key thing for intelligence.
[0] https://en.m.wikipedia.org/wiki/Encephalization_quotient
If beating the strongest human Go players with a self-trained ML model was detonating a fission bomb, pervasive automation with zero accountability is a nuclear winter.
So like a human being?
Humans are expensive and require individual training. Models scale. In the time it would take a human to make a mistake, a broadly-deployed model might make millions of impactful mistakes.
Even with all this analogy, we don't insist medicine should have known mechanism of action, although we do prefer it. So I think we will regulate models with recalls, testing standards, monitoring etc, but won't insist on understanding.
By the way, don't we already have widely deployed models, such as PhotoDNA, which supposedly removes millions of images a year to filter child pornography? I wonder how it was evaluated to be suitable for deployment.
I regularly go to the grocery store and purchase things produced by human from around the country and around the globe. In the process of producing things and getting to the grocery isles, I'm sure there are always fuck-up on various people's parts and there are other people who compensate for them. That behavior is far what AI can do at present.
That isn't saying AI isn't impressive in many ways. But it's a brute-force simulation of what people do and it's fragility demonstrates this.