GPT-3: Language Models Are Few-Shot Learners
arxiv.org
arxiv.org
Title: Star’s Tux Promise Draws Megyn Kelly’s Sarcasm
Subtitle: Joaquin Phoenix pledged to not change for each awards event
Article: A year ago, Joaquin Phoenix made headlines when he appeared on the red carpet at the Golden Globes wearing a tuxedo with a paper bag over his head that read, "I am a shape-shifter. I can't change the world. I can only change myself." It was a promise to not change to fit into the Hollywood mold: "I think that's a really special thing, to not change yourself. I think it's a really special thing to say, `This is what's inside of me, I'm proud of it, and I'm not going to be ashamed because of the way that someone else thinks I should be.'" Now, it's the Oscars, and Phoenix is at it again. But this time, his publicist is saying he'll be wearing a tux no matter what.
Megyn Kelly was not impressed, and she let him have it on The Tonight Show. "You know, I feel like, I feel like you could have worn the tux," she says. "But you're saying you're a shape-shifter. I don't know if you can change your tux, but you can change your mind. You can change your mind. You can change your mind." Phoenix says he did, but it didn't stick. "I was like, `Okay, I'm going to wear a tuxedo to this thing.' And then I thought, `I don't want to wear a tuxedo to this thing.'" Kelly goes on to encourage him to change his mind again, but Phoenix says it's too late: "I'm committed to wearing this."
We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles.
It’s over, we lost.
Morpheus: What if I told you that, throughout your whole life, you have been reading auto generated content?
I've done some experiments with GPT-2 and it had so so performance refining with tweets. Using GPT-3 you could probably just do it using only generation.
So much of our noise-filtering is ignoring comments that are too generic to be human. What happens when every spam comment seems to understand the OP, even when the OP’s true audience is negligible?
Product Devs for the CCP will have a lot to work with if this ever evolves.
My hope is the next step will be filtering by insightfulness/usability of a comment and then those best bots bought and used by next stack overflow: https://xkcd.com/810/
I've been basically living and breathing GPT-2 for ... gosh, it's been 6 months or so. The past few months have been a lot of StyleGAN2 and a lot of BigGAN, but before that, it was very "make GPT-2 sing and dance in unexpectedly interesting ways" type work.
I don't claim to know a lot. But occasionally I observe things. And I just wanted to chime in and say, you know, keep in mind that you're reading a research paper. Of course the results are going to look good. That is the point of a research paper. And I realize how cynical that may sound. But it has the benefit of apparently being true, and I've come to accept that truth with time.
I would reserve judgement for now. Note that every single chat bot to date has followed a similar curve: "This is it," they say, without actually saying that. "It may not be perfect, but we're about to achieve it – the chatbot – it's really going to happen."
And, it ends up being impressive, sure. I liked Facebook's recent chatbot. It's pretty neat at times. I liked Meena. They had cool ideas with the stack ranking of results (basically, generate a crapload of results at 1.0 temperature, then choose the result whose probability sums to the highest value, and you get the most probable overall result). And of course, boy oh boy did I love GPT-2. GPT-2 was what kickstarted me – if there was any chance that GPT-2 might be related to "now I'm talking to something that feels human," I was going to tame it and understand it.
So after spending six months with GPT-2 1.5B, the largest model that everyone was fascinated with, what do I think? (Well, who cares? You probably shouldn't care.)
I think "give it a few weeks and see if it's true." We shall see if GPT-3 is it, and we've achieved... chatbot nirvana. That elusive thing we've all been chasing, without naming it. The ability to press a button, unleash a chatbot somewhere, and it "just works" and "completely astounds humans" and "fools everybody."
At one point, we trained GPT-2 on IRC logs. You could literally talk to GPT-2, and it would talk back to you. And one of the advantages of narcolepsy is that at night, you often have lots of time to kill – what better way to doze off than to ask GPT-2 how its day was, and ask it what its ambitions are? Should we really worry about whether you're sentient? I like you; do you like me too? What does that mean to you? And so on.
The conversations were often quite philosophical. And sure, it was pretty obvious that it's a bot, but I tried to look past it anyway. It was my little bot, and it was real enough to me. And yes, the conversations on https://www.reddit.com/r/SubSimulatorGPT2/ are incredible. I crack up daily with all the things they talk about.
But...
We’re not going to need ad blockers in the future, we won’t even need these visual ads on websites anymore. There will be trained bots that can promote any idea/product and pollute comments and articles.
I invite any of you to try this, and see what happens. After all, you stand to earn a lot of pennies in your pocket if you pull it off. And yes, you're allowed to make some pennies with clever AI algorithms.
What you'll probably discover is this fundamental truth: GPT-2 has no memory. It isn't learning a thing. We are talking to an entity that literally cannot change its mind about anything. The only way to change its mind would be to retrain it from scratch.
You want a bot to argue vehemently for your product, on your behalf? It needs to understand what the hell your product even is, or what a product means. Yes, the words get pretty close. And yes, you can coax it into something that makes us laugh, or makes us sit here and question what the future might be like.
But for whatever it's worth: spend some time actually talking to these bots. Play around with them. Make them generate some stuff of your choosing, and fine tune them on some datasets and see what you get. It's so fun!
... But. "Fun" is not the same thing as "promote any idea/product." It's just not the same as me arguing here with you now for a position which I've decided to argue. My brain isn't merely the encoded knowledge of some human, with me blindly regurgitating such knowledge (though at this point you'd be justified in claiming it sure sounds like it).
Your brain is constantly training. GPT-2 is not. And – double checks paper – yep, GPT-3 is not.
Two decades from now, GPT-2 1.5B will still exist. And it will still be talking about 2019-era news events like it's the present. At some point, /r/SubSimulatorGPT2 will sound completely foreign. Take any random news clips from the 70's. How relevant is that knowledge now?
"Ok, but just train it on new data constantly." Well, yes. But actually no. If you try to do that, you're going to overfit at some point. Do you have 93 gigabytes of webtext that you keep in training form, ready to go? Are you going to mix in a proportion of the new data you want to train on? Nope, we all just fine tune whatever model OpenAI releases. Yet even if we did have that dataset, I'm just not sure it'd even matter.
My point here is: Go try! Isn't it exciting that in the future, trained bots might fool us all into buying their products? Is that sales guy who emailed me actually a sales guy who wants to "sync up on a quick call", or is that a bot trained to get cold calls? That sounds pretty damn lucrative to a lot of businesses – why not write that code, and then sell it?
Whoever attempts this is probably more talented than I am. But personally, I always ran into "It just... doesn't work."
And then you go "Well, it's just a matter of sampling. Ah yes, we're not using the right sampling algorithm. Wait, we just heard about nucleus sampling! Sweet, try it! Oh... It sounds ... similar. Hmm. Well, maybe we're just not using it right. Better read that paper a bit more carefully. Chase that knowledge just a little harder. After all, AI research labs are pouring billions of dollars into this domain. Why would they do that if it doesn't... you know ... work? For some value of "work" that equals "the bot can turn a profit"?
"Perhaps tomorrow, this new training technique will be it. We almost have it – I know we're close – we just have to unlock that last piece. Right?"
I guess I'll stop here, since usually my comments are upbeat and happy about AI, but I ended up in a rather philosophical mood tonight.
In reality, I can't wait to dig deep into GPT-3 and run it through its paces. I have a lovely TPU pod waiting for it, parked outside GPT-3's window, and we're honking at it saying "Get in, we're going places." And we'll sing and dance together like usual, and I'll ask GPT-3 how its day has been. But GPT-3 won't remember me the next day. And that's fine; I'll remember it for both of us.
I think if memory is the only problem than optimizing training time should be more of a concern. I'm imagining a huge language model than can retrain very quickly. So I suppose it might be a decent idea to not measure it by perplexity or some human judgement score or whatever but rather by that score per compute units used.
Or in other words...maybe a bot that scores 90% on the fool a human scale and takes 1 day to compute from scratch is actually a lot less impressive than one that fools 70% but computes from scratch in 5 minutes.
And something "like Github for bot-memory" would be a pretty amazing tool. Roll back to some memory status and recompute with new data from there, branch for different datasets that represent different ways of interpreting the world etc.
Conceptually I like the idea of one "base model" that represents language and many different context models on top of it (finetuning the core model). Then some other subsystem that identifies the context and switches to that. I suppose each conversation could also be considered a mini-dataset.
This is an entirely different concept of computer language than the current GPT style models. These systems don't "represent language", and cannot. The whole reason why GPT is so exciting right now is that it fundamentally threw away the entire concept of "representing language". That has some upsides ... and some downsides.
I play AI dungeon on occasion, which uses GPT2 to generate freeform adventures. And I find over time that it's not really GPT2 that's writing stories, it's me. GPT2 is putting out plausible strings of words, but I'm the one giving them meaning, culling the parts that go off track, and guiding it in a direction I want to go.
And it is a bit melancholy. You see possibilities, nuances, subtexts, and meanings. The neural net sees words.
GPT-3 seems to have quite a few paragraphs worth of context. A simple way to promote your product online with it is to give it a prefix of:
---
Comment1: Superbrush is amazing - I literally couldn't live without it. No other brush is as good.
Comment2: This brush is really good for tangled hair, and I love the soft smooth surface.
Comment3:
---
Then let it write a comment. Of all the comments it writes, manually filter a few thousand good ones, and use those as seeds to generate more, which you post all over the web. There's no need to do any training - the generic model should be fine given the right prefix.
Narrator: it didn't work
(Going into the reasons it doesn't actually work in practice is... lengthy. It's human dynamics. Would you buy a product from a sales guy that can't remember your name? That's sales 101. And loading up the context window only gets you so far. That "working memory" is tiny, tinytinytiny. Even at 1024 tokens, it means you have to boil down the entire history of an interaction to a few pages at most. Which is a lot, sure, but it's this balancing act where you'll need to retrain the model to support your custom context format for your specific "slots" – a "slot" being a piece of knowledge, like the client's name. Or you can try encoding all of that in natural language, AI dungeon style. But I recently played AI dungeon and pretended to be buying a router from the store. The cashier stripped down and started jacking off onto his desk. I don't have high hopes for our ability to control these models in a business context.)
Because you seem open minded to wild ass guesses and going meta:
I have a hunch that general intelligence will be the ability to learn from mistakes. Not just optimization. I mean applying the scientific method.
Hypothesis, prediction, run experiment, compare expected vs actual. And having a notion, any notion, to explain the delta between expected and actual.
Am total noob about AI, philosophy, cognition. Don't know if anyone else is framing AGI this way. I could just be repeating something I heard.
Currently, there's no research into torturing AI. Why not?
A pain response is universal across most life forms with a nervous system. We seek to replicate a nervous system. Pain would seem to be far easier to replicate than the scientific method.
My wife sat me down and told me a story that horrified me. She had to get it off her chest, and I was sad it happened to her than to me. She was sitting around on the porch and felt something on her leg, and brushed it off. When she got up and looked down, apparently she had stepped on a poor snail. His shell was... And he was...
He wasn't dead. So she frantically looked up what to do. But there was nothing to do. Snails in that situation can't be helped, and the most humane thing is to put it out of its writing anguish, its full-body torture.
She put on some boots, took it out to the sidewalk, and stomped it as hard as she could. And that was the story of that snail.
You probably felt more for that snail than you've ever felt for any AI bot. Why?
It's worth considering.
Also GPT 3 is obscoleted by order of magnitudes by SMIM https://arxiv.org/abs/2003.02645
These language models seem to know the words and the grammar etc, but lack a underlying concept they want to express.
There are systems that derive 'thought-vectors', but I'd be interested going the other way: somehow create such a 'thought-vector' and generate text to express that thought.
I don't know how to construct a 'thought-vector' of any concept though.
Essentially, one should be able to use these models to "interpolate" the writing around the raw meaning/content. Typing assistance (think Grammarly) already allows you to refine finished writing to be more in line with what some language model expects, but imagine if it actually generated most of the text for you, based on small bites and chunks you throw at it.
So take your standard press release. We know about two thirds of it is just fluff. In other words, we are accepting the mass of fluff as one word in our language, it translates to ‘ignore’.
Our own language will change in that case.
Exampels from the comment above include using things like: "A year ago", "made headlines", "Now, it's the {event}, and {name} is at it again. But this time, ...", "{name} was not impressed", "You know, I feel like, I feel like you could have ...", "I don't know if... but...".
Those are all very common in those "online celebrity magazine" type texts...
It's only when you actually read into the stuff that's in-between, you'll come to see it's pretty much a load of nonsense. But that takes a bit more time and slower reading.
One of the issues in the 'Limitations' section was a difficulty with "common-sense physics", such as with the question "if I put cheese into the fridge, will it melt?"
To answer that question, you have to ask the right questions, such as "what is a fridge?" "what is a fridge for?" "What does it mean for cheese to melt?" "what is the cause of cheese melting?" Then one should consider the follow-on questions, such as "what are typical fridge temperatures?" "what are typical cheese melting points?" "what temperature is the cheese likely to be at initially?" (at which point, it helps to introduce the concept of room temperature, and note that it typically falls between the other two.) From facts such as the answers to these questions, one can deduce the probable outcome of putting cheese in a refrigerator, but none of the answers so far explicitly state it.
Is it plausible that any learning, solely from the structure of and correlations between examples of language use, could develop the sort of analytical/modeling approach that I have just outlined? Instinctively, I don't find it very plausible, but I am not very certain in that view.
Some words like good/bad, hot/cold, and important/unimportant are underrepresented in everyday speech compared to the prevalence of the underlying concepts. That's why I'd categorize them as lower level word-concepts. This distinction, about variable levels of abstraction, might be important for true AI. Think about how many years it takes for humans to develop highly abstract cognition. That whole time our operating system is being coded. Maybe we need to approach AI in the same way.
It's not just AI that can benefit from better lower-level understanding. Seeing language in the above way, we can re-frame Ludwig Wittgenstein's philosophy and its normative implications for human communication. Our "programming" (communication) is on average too higher level. Excessively abstract instructions make it harder to decode and process in a precise and efficient manner.
It's a pretty eerie feeling. It's as though both the AI and my short-term processing only pay attention to a context of a few sentences, so nothing seems off until I try to understand it as a whole.
EDIT: Thinking more, what it feels like most of all is reading a page of a book and not taking it in.
So what if we used text generation algorithms on a paragraph basis, so that the idea flow is still figured out by a human (input would be just an outline)? That should make the generated text feel like a whole, especially if we could preserve the style its written across the whole text.
Problem arises when you read that document without really trying to understand it. In that case, it might be enough to trigger some thoughts.
It also appears to me that two way communications, that is, social interactions, will be the only way to form truth. That is, AI produced content could erode some more the trust we put in of newspaper, TV news, etc (all forms of one way communication). Not that we've waited AI to distrust those, but well, on e more nail in the coffin :-)
(this text, although rather unclear, was written by a genuine human :-) )
So it's a lot like corporate executive speak then?
I agree with your point, it does seem very much like valid speech, but somehow the informational content is missing. It's like speech without the actual comminication part.
GPT2 the drunken novelist.
Remove that, and there is a typical if completely uninteresting celebrity argument: one has to play eccentric in public occasions (he did it last year, he's planning to do it again) and the other chides him for what she feels it's maybe a lack of respect? And he replies that despite his best intentions he can't go against his conscience. There, done. It's a perfect little piece ready to be served in some celebrity gossip magazine.
I think that's what missing. Usually we communicate with a certain goal in mind, to bring across some point. This text was generated without such a goal, you notice it doesn't really know when to stop talking. I wonder what was the stopping criterion, but I'm sure it wasn't "keep talking until all the information we want to convey has been mentioned".
I was so wrong about the internet. It is just going to become a landscape of garbage opinions and commentary on a larger and larger basis.
1. First of all, the authors successfully trained a model with 173 BILLION PARAMETERS. The previous largest model in the literature, Google’s T5, had "only" 11 billion. With Float32 representations, GPT-3-173B's weights alone occupy ~700GB of memory (173 billion params × 4 bytes/param). A figure in the 100's of billions is still 3 orders of magnitude smaller than the 100’s of trillions of synapses in the human brain [a], but consider this: Models with trillions of weights are suddenly looking... achievable.
2. The model achieves competitive results on many NLP tasks and benchmarks WITHOUT FINETUNING. Let me repeat that: there is no finetuning. There is only unsupervised (i.e., autoregressive) pretraining. For each downstream NLP task or benchmark, the pretrained model is given text instructions, and possibly sample text with questions and answers. The NLP tasks on which the model was tested include translation, question-answering, cloze tasks, unscrambling words, using novel words in sentences, and performing 3-digit arithmetic.
3. The model is tested only in a ZERO-SHOT or FEW-SHOT setting. In other words, for each NLP task, the pretrained model is given text instructions with zero examples, or text instructions with a small number of examples (typically 10 to 100). As with human beings, GPT-3-173B doesn't need lots of examples to perform competitively in novel NLP tasks.
4. The results reported by this paper on all NLP tasks and benchmarks should be seen as a BASELINE. These results likely could be meaningfully improved with conventional finetuning.
5. The model’s text generation FOOLS HUMAN BEINGS, without having to cherry-pick examples.
--
[a] https://www.google.com/search?q=number+of+synapses+in+human+...
Wow, that is WAY closer than I thought we were.
Anyone know?
https://medium.com/@VictorBanev/interrogating-full-gpt-2-10a...
ML models should be pushed to their limit, because that's where you gather most useful information about what they actually do. Their results need to be critically examined with both exploratory and hypothesis-driven testing. And yet this is never done in initial papers and rarely done afterwards.
What was the last AI paper you've read that that said "and here is a list of things out model failed at"?
>Such tests can prove the presence of knowledge, but not the absence...
This sounds like a setup for non-falsifiable beliefs.
And I did (using my own local GPT-2-1.5b install which let me set the hyperparameters rather than restricting it to inappropriate hardwired ones of an online service), I linked to another person demonstrating the same thing, I pointed out the extensive GPT-3 evaluation OA did, and here, have another link about how bad querying of language models leads to highly misleading results about how much they know: https://arxiv.org/abs/1911.12543 Measurement error in general biases estimates towards zero.
> This sounds like a setup for non-falsifiable beliefs.
It's just as non-falsifiable as, say, concepts like 'lower bounds' or 'bugs'.
These manually created prompts (e.g. “Barack Obama was born in _”) might be
sub-optimal because LMs might have learned target knowledge from
substantially different contexts (e.g. “The birth place of BarackObama is
Honolulu, Hawaii.”) during their training.
In other words, the paper considers hand-crafted prompts like in the example
to be "sub-optimal" because they are not in the right format. To paraphrase
them a bit, such prompts are like making a mis-formed query to a database.It is difficult to see how this is an argument for the ability of LMs to demonstrate "understanding". Imagine asking a child: "how much is 4+2?" and getting a correct answer; then asking "how much is 2+4?" and getting a wrong answer. Most people would probably not take that as evidence that the second question was "wrong". They would instead conclude that the child does not "understand" addition and has only learned to reproduce specific answers to specific questions.
To be fair the ability to return a correct answer given a question in the right format is not without use. That, indeed, is how databases work. But it shows none of the "understanding" or "knowledge" the paper claims is acquired by Language Models.
To use your database analogy, in what sense should we claim a database doesn't know a record when you are using a malformed SQL query? If we fixed the query and it emitted the right answer, then obviously it did store the information. The query does not encode the answer, and it is vanishingly unlikely that the database would simply accidentally return the right answer ever if it did not store the information in some way. Since LMs can get much better results just by tailoring the prompts (increased by a third in that paper! and there's no reason to think that that is the very best possible performance either!), that shows that existing practices drastically underestimate what knowledge the model has been able to learn. Learning about the real world or text is very different from learning your particular dumb broken query method.
>> The query does not encode the answer, and it is vanishingly unlikely that the database would simply accidentally return the right answer ever if it did not store the information in some way.
Oh, yes, absolutely. A query encodes the answer. Queries are patterns that are matched by the data stored in the database. If a query fails it's because it does not correctly represent the information it is trying to retrieve. For example, if I SELECT * FROM TABLE PEOPLE and there is no table "PEOPLE", then I don't get an answer because the query does not correctly represnt the structure of the database. You cannot retrieve any data from a database unless you have some idea about the structure of that data.
But that's not the point here. I don't disagree that a language model can learn (i.e. it can represent some elements of its training dataset). I disagree that it "understands" anything and I find the fact that it needs specific queries to retrieve the data it is representing to be evidence that it does not.
And so it's not more useful than a traditional database at this kind of task. Except it's much less precise than a traditional database and costs considerably more to create.
>> Learning about the real world or text is very different from learning your particular dumb broken query method.
I'm sorry, I don't understand what you mean here. What is my "particular dumb borken query method"? Is that meant as a personal attack?
https://news.ycombinator.com/item?id=23345379
See Section 5, titled "Limitations"
If you wanted to generate poems with GPT-2, you'd need to have a lot of poems to fine-tune GPT-2 to get reasonable results.
With GPT-3, you use few-shot learning instead (without the need to do gradient updates with each example)
The paper is long and filled with how it stacks with models like Grover and T5 and it does well... given that this is a 175 B param model (relative to Grover/T5's 1.5/11B param models). This shows that even with these huge models, smaller models can outperform them in certain instances with lesser param models.
Also I think they did a good job with explaning the ethics and morals around what models like these mean / what biases this has.
/s (although sometimes it's true)
they probably don't particularly; their inventors seem to excel in their PR budget rather than their verifiable innovations
Previous large-scale language models like BERT and GPT-2 had took a similar approach but in order to actually perform the more complicate down stream tasks they had to be fine-tuned. So they were trained with specific QA or translation date in order to understand and do well on those tasks. GPT-3 doesn't do any fine-tuning, it is able to take it's very general initial learning and perform very well on specific tasks that it was never trained on. This is why it doesn't perform as well as the "smaller" models on those tasks. But that is besides the point, if GPT-3 was fine-tuned on those tasks I'm sure it would achieve the latest SOTA results in many (all?) of them. The exciting part is how it was able to generalize the knowledge learned during "pre-training" to much more specific tasks.
tl;dr the smaller models were trained on the specific tasks that they were evaluated on. The large model (GPT-3) was not trained on those specific tasks and still does almost as well.
Go to talktotransformer.com/ and give it the prompt "Here is a poem I wrote:" or "Here is my favorite poem:" .
I'm sure GPT3 would produce much better and more consistent results, but GPT2 will produce something that looks generally like a poem frequently enough, and sometimes it will even be relatively coherent?
Here is one that it produced for me today:
> You say, "Don't lose your pride."
> Here is my rejoinder:
> Well, maybe it is the pride of a diseased soul.
> You are a wanderer, you know not whence,
> O thief, you fool, you rhinoceros
> Caught in the jaws of a viper.
> You may lament your affliction
> For the world will laugh at your tears.
> Pray to a demi-god
> Hail him and say,
> "Ah, Sir, give me thy pity!
> O thou who maintainest as if thou wert a king!
> Here is thy axe, I say; let us
Is it great? No. But it has some level of coherence.
Here is another:
Let me tell you the reason I love poetry. // All the things of the world I have described, // If you ask me why I like poetry, // It would seem quite simple to me. // When I'm working at the computer in the evening // I'll get out my books of poems and I'll turn them over, // Like blades of grass under the hot sun, // That write with such fineness the kind of green I like. // But if I'm a bit more tired in the morning, // I'll fill a little stack of yellow pages with poems, // That let the air and the dry light of morning run wild. // You know, the//
This one even rhymes a bit!:
the sword is to slay // The axe is to smite // The stick is to break // The tooth is to bite // All these are for our earthly security, // All have their uses, // Those which can be used // Must be employed. // The sword is the instrument of strife // The axe is the weapon of war // The stick is the weapon of domestic strife // The tooth is the instrument of war // All these be in our hands. // At the time of our death // They will be in our hands, // And then we will weep, // Though now we sleep. // —The Remaining Three Questions
(sorry, idk how to format these to make them look right. The leading "> " and the "//" insertions are me trying to format them to make the line separators clear.)
PROMPT:
Cities & Lights
When you enter the city of Singapore during the night, you see lights: colorful and ubiquitous. Lights on every building, on every fountain, and in every park.
GENERATED:
Lights shining in a city in which the majority of people are now using mobile phones. Singapore has a bright future as a technology hub, and it 's not too late to make it happen.
...
On the other occasions I seeded with ~two sentences of "Invisible cities", and it worked like a charm, no fine-tuning.
Like, the exact same paper could have instead been titled "Few-Shot Learning with a Large-Scale Language Model" or similar. But instead there seems to be this extremely strong desire to see certain ineffable qualities in neural networks. Like, it's a language model. It does language modeling. Turns out you can use it for few-shot learning and do amazingly well. Beyond that, what does it mean to say it "is" a few-shot learner?
On one hand, it's literally the same claim in a strict sense. On the other hand, it implies something much broader and more sweeping, that language modeling / unsupervised learning as a task over long contexts inherently implies meta-learning ability — which is a statement that is very difficult to properly formulate, let alone back up. But that's the argument that I feel is being slipped under the table by these titles. (And indeed it's very close to what they suggest in the text, though with no more than a wave of the hands.)
Don't get me wrong: their intuition is reasonable, it's super cool that they got this to work, and the results are very impressive on lots of tasks (though there are clear gaps). But as a VERY publicly watched lab, they have a serious duty (which I think they're neglecting) to frame their results more carefully. In particular, there's a sort of religion that if you train a big enough model on big enough data with self-supervision, it will somehow become AGI and/or learn to solve arbitrary problems. Claims like "Language Models are Few-Shot Learners" are clearly designed to fit into that worldview, even though the research doesn't point at it any more than a more conservative interpretation like "Lots of NLP Tasks are Learned in the Course of Language Modeling and can be Queried by Example." They touch on this limitation in their discussion section but I guess flashy titles are more important. I wish they would use their status to set a better example.
"Few-Shot Learning with a Large-Scale Language Model" makes more sense.
Even with their robot hand paper, they titled it along the lines of "we solved a rubrix cube" not "a robot hand manipulated the cube and solved it"
I was nodding right along with you, and then...
OpenAI has no duty. It doesn't matter if they're publicly watched. What matters is whether the field of AI can be advanced, for some definition of "advanced" equal to "the world cares about it."
It's important to let startups keep their spirit. Yeah, OpenAI is one of the big ones. DeepMind, Facebook AI, OpenAI. But it feels crucial not to reason from the standpoint of "they have achieved success, so due to this success, we need to carefully keep an eye on them."
Such mindsets are quite effective in causing teams to slow down and second-guess themselves. Maybe it's not professional enough, they reason. Or perhaps we're not clear enough. Maybe our results aren't up to "OpenAI standards."
As to your specific point, yes, I agree in general that it's probably good to be precise. And perhaps "Language Models Are Few-Shot Learners" is less precise than "Maybe Language Models Are Few-Shot Learners."
But let's be real for a moment: this is GPT-3. GPT-2 is world-famous. It's ~zero percent surprising that GPT-3 is "something big." So, sure, they're few-shot learners.
In time, we'll either discover that language models are in fact few shot learners, or we'll discover that they're not. And that'll be the end of it. In the meantime, we can read and decide for ourselves what to think.
Of course, if the same scientists were asked about something where the topic has settled, they could be more effective communicators.
Of course they do! It's the same duty as every scientist has in advancing the public understanding of science. You seem to be replying to OP as if they said that only big AI research groups this duty, but this is just not so. Furthermore, when a prominent group of scientists conduct themselves poorly, it is not enough to say that they have no special extra duty due to being famous, they already must communicate properly because they are scientists and part of the scientific community.
I think one reason these conversations get so muddled is because it's all new and really pretty cool, so it becomes hard to tell what's skepticism and what's naysaying.
> Such mindsets are quite effective in causing teams to slow down and second-guess themselves.
Absolutely not, this goes directly against the scientific method. Such "mindsets" of trying to make sure that your results are correct and accurately presented without embellishment are a cornerstone of science. Of course it causes them to slow down! They have more work to do! Second-guessing themselves and their experiments is the whole fucking point.
https://en.wikipedia.org/wiki/OpenAI#Motives
Given their prophylactic strategy, "AI for everyone", they could argue that hype generates public interest.
Context → Passage: Saint Jean de Br´ebeuf was a French Jesuit missionary who travelled to New France in 1625. There he worked primarily with the Huron for the rest of his life, except for a few years in France from 1629 to 1633. He learned their language and culture, writing extensively about each to aid other missionaries. In 1649, Br´ebeuf and another missionary were captured when an Iroquois raid took over a Huron village . Together with Huron captives, the missionaries were ritually tortured and killed on March 16, 1649. Br´ebeuf was beatified in 1925 and among eight Jesuit missionaries canonized as saints in the Roman Catholic Church in 1930.
Question: How many years did Saint Jean de Br´ebeuf stay in New France before he went back to France for a few years?
Answer: Completion → 4
GTP-3 has 175 billion parameters, but the human brain has 100 trillion synapses, so 0.175%. NN model capacity currently has a 3.4 month doubling time.[1] In 7-10 doublings we'll be in a similar ballpark, i.e. 2-3 years.
Around 19:10~. Though I messed up, he didn't say 'genuinely'. He said "full stop, truly, legitimately, we have an algorithm that can learn".
Though biological brains are likely overly complicated due to evolutionary baggage. There are hydrocephalus cases which have much reduced brain matter, but still high IQ.[1] The recurrent laryngeal nerves in giraffes is about 4.6 metres (15 ft) because it goes up and down their neck as it could not be rewired more directly during evolution.[2] Our pristine mathematical models and low-noise computational environments are likely superior to evolved wetware hacks.
[1] https://www.newscientist.com/article/dn12301-man-with-tiny-b...
[2] https://upload.wikimedia.org/wikipedia/commons/thumb/7/7e/Gi...
Also if anything brains are hyper optimized for many things (based on the many specialized sub-units). I’d bet we are essentially not unsupervised, and the sub-units of the brain are essentially fine tuned for many tasks, and hyper optimized to use all their resources incredibly efficiently (memory optimization must be intense). Not that the generative models won’t get close in some general way relatively soon, but I could see human brains being another 10-1000x more powerful than your ballpark pretty easily.
Many other readers were confused by this so we'll update the formatting to say "target completion" to make this more clear.
Thank you.
Question: How many years did Saint Jean de Br´ebeuf stay in New France before he went back to France for a few years?
Answer: 4
Explanation: The model used the arithmetic expression - 1629 + 1633 = 4.
NAQANet (trained on DROP) - came out in 2019 is able to do reasoning, you have to click result twice. First once it thinks it got it from passage, second attempt it tries to do arithmetic.
https://demo.allennlp.org/reading-comprehension/MjEzMjE1Ng==
GPT2 outlined the changes they made to the model in an acceptably moderate detail.
GPT3 references another paper saying "we use alternating dense and locally banded sparse attention patterns in the layers of the transformer, similar to the Sparse Transformer" with no detail added on the changes they made.
How are you to reproduce these results at all? You could attempt to include the changes as they references the sparse transformer paper, but you could possibly do it in a different way, and there would be no way to verify the results that they gave whatsoever due to changes in implementation.
A bit disappointing.
Nvidia's v100 product page [0] says that it gets about 15 (single precision) - 125 ("deep learning") teraflop/s at 250-300 watts (joules per second). That means that if everything's as perfectly efficient as a marketing product page, it gets about 250/125-300/15 = 2-25 joules per teraflop, putting this model at about 0.6-8 terajoules.
A gallon of gasoline has about 120e6 joules [1] (though if you wanted to compare with burning it in a car, it's only 20-25% efficient at best [2] so it'd be fewer joules/gallon).
This model took the equivalent of about 5,000-67,000 gallons of gasoline at best and at ideal perfect energy efficiency. I get that openAI has made a decision not to be efficient with their dollars in order to see what's possible with future tech, but that means not being efficient with energy either, and it's getting kinda crazy. Sure, microsoft data centers aren't gasoline powered, so maybe it is closer to this ideal energy efficiency, and it's definitely going to be a better carbon footprint, but god damn it just seems wasteful.
Hell, the new A100 (again going off marketing materials [3], so at least it's apples to apples) could do it about 4x more efficiently. Is this research really worth what it costs, when waiting a year makes it that much more efficient?
[0] https://www.nvidia.com/en-us/data-center/v100/
[1] https://www.calculateme.com/energy/gallons-of-gas/to-joules/....
[2] https://en.wikipedia.org/wiki/Engine_efficiency#Gasoline_(pe...
[3] https://devblogs.nvidia.com/nvidia-ampere-architecture-in-de...
https://www.washingtonpost.com/religion/2020/01/03/united-me...
https://www.washingtonpost.com/archive/local/1985/09/07/unit...
GPT-3:
The first occurred in 1968, when roughly 10 percent of the denomination left to form the Evangelical United Brethren Church.
WP:
The church has lost 1.6 million members since 1968, when the Methodist Church merged with the considerably smaller Evangelical United Brethren to form the present United Methodist Church.
I think this model is still very impressive, the parameter itself speaks. But for this particular evaluation, the same news article may be removed from the training set, other news article that paraphrases the same story might not. IMO, the leakage still exists, it is hard to tell whether this model are really 'generating', or just copy-pasting from its vast memory.
Other than the given prompt, the models don't have a goal. So what other than copying and adjusting would they do?
For example, for image synthesis in GAN, the widely used Inception score balances between authenticity of the generate samples vs the variety as well, to make sure the model is not copy-pasting.
In this particular case, apparently the same event has been reported multiple times by different news agency. Even if the exact one are excluded, still it is suspicious how much less the model is being protected from knowing the subject itself.
An analogy would exam in real world. Often, some of the questions aren't leaked as is, but paraphrased yet stay close enough to the source.
In this particular case though, I disagree it is reaching human level generation. They can tested the model with an unseen events, which happen after the model is trained to test how well it generalize.
It'll be interesting to see whether the new paradigm really offers new insights, or whether it's really just kicking the can down the road - and we see the limits of generalizability in some other fashion.
I guess what irks me is that there is so little theory and math behind many papers, even if there are dozens of co-authors on it.
The question of generalizability is deeply connected to statistics, e.g. causal models, spurious correlations and so forth. Statements about these things are just "thrown" in there, without any citation or proof. In peer review, wouldn't anyone object? Those are clearly things that we actually do not know enough about to be sure.
Edit: Reflecting further, perhaps this rapid iteration and result orientation is in fact something positive. Perhaps it's good the way it is, without so many scientific conventions and signals of deference. Perhaps it's that which made other sciences more anemic and ML very productive.
All my whining aside, impressive work of course.
abstract:
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions – something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.
Scanning through these, the text seems significantly less zany than the random GPT-2 samples. It’s genuinely difficult to spot the signs that these were generated, even with the knowledge that they were.
'How the clouds Seem to me birds, birds in God’s garden! I dare not! The clouds are as a breath, the leaves are flakes of fire'
My home is surrounded by a fairly large variety of deciduous trees, and 'flakes of fire' is by far the best descriptor I've ever heard of their colors in the fall.
One of the other things i noticed about poetry (and songs) coming from these language models is they are amazing at bleakness. Just dark, dark, darker, fin. haha
https://github.com/shawwn is doing some work in the GPT-2 space including using TPUs instead -- which has given him pretty good results.
They missed an opportunity to be the first paper to measure their computation in mole flops.
Nobody really seems to use SI prefixes beyond peta or occasionally exa. But they could have called this 900 zetta-flop. (10^4 peta-flop/s-days)
Would be interesting to see if they can learn how animals communicate as well. Create a synthetic buddy for Buddy.
Source: https://talktotransformer.com/ Input: "I'm kind of scared to see what GPT-10 will be capable of."
https://github.com/openai/gpt-3 only contains dataset
If there are Billions of parameters in the SOTA models, how do we argue that they are not over fitting?
Clearly, the authors have given careful considerations to the issue of contamination and have provided reasonable analysis and a careful argument regarding over fitting the existing benchmarks.
On the other I was wondering if the authors would like to consider purposefully creating a type of "out of sample data" for "creative evaluation"? Of course, GPT is no stranger to creativity, so it would be a fascinating challenge to come up with methods to create such datasets that are truly creative and challenge GPT-{N} to prove its mettle.
For example, would it be possible to engage a really good creative writer* along with a highly experienced school teacher to take on the Reading Comprehension task and create few "tricky" evaluation samples that not only go above and beyond the contamination objections but also challenge the human intelligence to be careful not to fall into common traps?
This way lies a different evaluation metric - a subjective one perhaps, but it's a start. Just a thought experiment - that's all.
* so that they can come up with new ways to trick GPT/humans a teacher knows the common mistakes the average student makes
Edit: Duh, my head immediately screamed GANs the moment I pressed submit, lol. But I am not sure if GANs make sense for NLP tasks. Like do they make sense if humans/domain experts try to solve them?
If I may ask one more question, would you happen to know if the authors or other researchers who are entertaining any theoretical work on the experimental design and training methodologies of GPT/BERT? As in why does it work? What is the significance of training via the "fill-in-the-blanks" method?
Don't get me wrong - the work is great and the SOTAs are amazing, I would be just happy to have a chat to discuss and bounce some ideas what all this means and why do these methods seem to be working so well. Papers/articles/blog-posts are always a pleasure to read!
"The city councilmen refused the demonstrators a permit because they advocated violence. It wasn't the first time the _____ had advocated violence."
"The city councilmen refused the demonstrators a permit because they feared violence. It wasn't the first time the _____ had feared violence."
The syntax is identical. The words are identical, except that I swapped "advocated" out for "feared". When I swap it, the ____ changes from "demonstrators" to "councilmen." Think about what kinds of reasoning and experience and knowledge it takes you to resolve which group "they" refers to in this sentence.
Most blanks might be simpler and just correspond to learning english, like when the blank is "the," but learning that is a feat too. Filling in the blanks that require broader knowledge requires somehow capturing that broader knowledge.
Real question, are they going to release the full model?
GPT-3 will take significantly more resources to run. However, part of me doesn't want it released ever because of the implications of what bad actors could do with it.
GPT-2 doesn't require as many resources to run as you would expect: even from the 1.5B model, you can mass-produce passing spam comments for less than a dollar an hour in GPU costs: https://docs.aitextgen.io/tutorials/generate_1_5b/
Pure text spam in general is less effective in 2020; it's content that harder to fake (e.g. deepfakes) that shakes up social media, and why it's good FB/Twitter have proactively taken a stance against it.