A demo of GPT-3's ability to understand long instructions
twitter.com
twitter.com
> A caveat to all of these: I use GPT-3 a lot, so I know the “golden path” of tasks it can do reliably. Had I asked it to write a sentence backwards or sum a list of numbers, it would fail every time. These are all softball questions in isolation.
I haven’t shown that GPT-3 can handle all coherent directions of this length, or even most directions that an untrained person would think to create. It’s just a demo that, if GPT-3 happens to be capable of your tasks separately, length per se is not a major issue.
That's kind of the whole deal of the attention mechanism in transformers and also partially why they replaced RNNs. You don't throw away any part of the original input as you construct your output. The downside is that unlike for a RNN, the total sequence length is fixed at training time and complexity grows with the square of it. But apart from computational cost, sequence length is not really an issue anymore for these models.
Funnily enough, that is exactly the kind of thing a human might do, as we too are terrible at following instructions precisely.
I also suspect it was confused by the fact the name was abbreviated but not misspelled, and it was only told explicitly to ensure names are not misspelled. Still an error though.
Someone one twitter suggested that possibly the algorithm interpreted that as "abbreviated" but not "mis-spelled".
Either way it's intriguing.
I will say, though, that it bugs me when people say it can’t be conscious because it sometimes says stupid things. In most cases, there are known tricks to suppress undesirable behaviors. More to the point, though, if we encountered a human being who gave confabulated answers to questions like “When was the Golden Gate Bridge transported for the second time across Egypt?”, we wouldn’t insist that this human is not conscious — we would just call them brain-damaged or mentally ill. I don’t think modern models are conscious, but as a logical possibility they could be conscious and still say very stupid things all the time.
A biological system self-modifies even when just associating memories. The phenomenon you describe is much more high level, we observe someone can't form new memories, but that doesn't mean their fundamental brain chemistry has changed.
But then I realized that somewhere on the Internet there inevitably is a message board where people play the "find me some shit on the internet" game, and there's some rabid subculture around it with zillions upon zillions of of examples, and it's in the Bing index, and all the nudging it would need is to emphasize that sort of thing in the corpus.
Very impressive stuff.
It's my opinion that these hyper-scaled transformers are actually a great deal less mysterious than seems to be in the zeitgeist, but for reasons that actually make me think there is a lot of headroom on capability: when the corpus is basically everything ever digitized like it is when a search or social network megacorp trains one, the only thing it could never do is something literally unprecedented on the Internet.
The mechanism can be good old `P(thing|internet)`, but if the KL-divergence is low enough, sampling from the modeled distribution can write something like Tristan und Isolde or paint something like the Mona Lisa.
Makes me double down on my prediction a week or so ago* of a Mid-Level AI Knowledge Work Apocalypse. In the next decade, AIs like this are going to do to office work what robotic mechanization did to the manufacturing sector.
All truth passes through three stages. First, it is ridiculed. Second, it is violently opposed. Third, it is accepted as being self-evident
At the end of the day, all the demos leave me with a feeling of disappointment. Current image synthesis models appear to be useless beyond doing experimental art for fun and novelty, chatbots still suck and copilot just (sometimes) replaces googling, but not developers or their education.
look up companies like uipath
The closest to it was perhaps that code-generating demo here a day or two ago - but who wants to be a 'GPT programmer' writing code as 'write a Python program that computes fizzbuzz replacing the arguments $fizz$ and $buzz$, ...' instead of just the 'actual' code? It just seems like a more clever AppleScript to me, pseudocode, and I don't think anybody's ever seriously pursued a flexible keyword pseudocode like language as a goal, it's just appeared as a demo of more general models?
Generating template/outline text I suppose? (Like that essay-writing helper here a few days ago.)
https://github.blog/2022-07-14-research-how-github-copilot-h...
As to what else future language models could power, based on my own use, I think fine tuned future language models could probably handle most customer support, accelerate the creation of most web content, accelerate quite a bit of paralegal grunt work, and power highly interactive game NPCs like in AI dungeon, another launched and paid product based on GTP-3.
To add, github copilot is a really clever autocomplete that makes some mundane tasks much quicker. Things which are too small for a library but are still fairly often used can be "typed" more quickly.
I’m assuming there are existing solutions: paid services or Python libraries undoubtably. But it is a lot easier for me to take 10 minutes of my time to put together a prompt and add the API endpoint to my Nodejs app. As far as cost goes, our needs are low-volume so it doesn’t really matter.
I see GPT3s great utility here as a replacement for Mechanic Turk type tasks. MT is a real headache to setup and manage programmatically. GPT3 is pretty simple once you integrate it into your system. And the fascinating thing is that whole new realms of functionality are opened up to product ideation.
I haven’t tried to do anything beyond low-volume tasks that are Mechanical Turkable, but with GPT4 I think we’ll see cost and performance drop to the point where these things can be done at scale. At that point, software engineers would be foolish not take make GPT4 just another standard library they have at their disposal — kind of like how jQuery opened up web development by providing a generic interface to DOM manipulation across all browsers.
It is just as hard for us to imagine what kind of development GPT4 will enable as it was for people to imagine an internet of web-apps before jQuery. But it certainly will.
In the simplified diagram below the network reaches the answer in the second layer at node "X", and reports derivations of it at 2 positions (obviously there are many more nodes and its a bit more complicated, see https://en.wikipedia.org/wiki/Transformer_(machine_learning_... as GPT3 is a transformer neural network)
O O. O
X.' 'O.
O 'O. 'X
O 'O.
O O 'X
O O
O O OIt's basically one of those markov chain bots, except with a very advanced statistical model behind it.
[*] Technically it's not words, but tokens. GPT tokenizes text to better compress the amount of data being fed, it's basically a big vocabulary list that compresses text into a list of ints, for example the string "hello world" would be converted to the list [31373, 995]. For this case it's an int per word, but less common words will not be compressed this well, with the worst case scenario being a token per letter.
I should also note that while the model's forward pass only works with one token at a time, the text generation is more advanced than that, there's multiple methods like beam search and top k-sampling, each with their own settings and tunables, but the gist of it is that during generation it'll try multiple combinations of token sequences and check which one is the most likely.
The limitation is memory, transformer networks are notoriously memory hungry, and IIRC GPT-3 grows quadratically with the number of tokens given, usually the limit is around 2048 tokens, or roughly 1000-2000 words.
It might be interesting to see how well it is able to modify the first output given some aspect of the final tasks.
1. Read through all steps carefully.
2. Do X
3. Do Y
(...)
99. As you have now read through the instructions, simply put your name in the top right corner of the first page.
Is it able to extract any kind of structural information? For example, you pass it the text of a movie script or children's story (where the descriptive language is simple) and it returns a structured summary of the content?
Summarize the following text:
I don’t really know what to say. It’s taken so long to get to this point, but here we are. Through all the trials and tribulations we’ve faced, it all comes down to this. You, me, and the unmistakable facts of our situation. This is all that remains: The truth. The truth is something we can’t escape, or at least you can’t — not anymore. Because the truth is that you have left my pineapple slices out of the refrigerator, and thus I will not be able to partake in their joyous, fruitful delights. How dare you. You scoundrel. You wicked, wicked thing.
Answer:
And the completion given: The text is about a person's anger at someone else for leaving pineapple slices out of the fridge.
It can, demonstrably, summarize text. The fact it sometimes makes mistakes for some texts doesn’t change that fact.Edit: Here’s an example summarizing 8,294 characters of Harry Potter fan-fiction: https://twitter.com/goodside/status/1561213457374011392?s=21...
How do they know that a new version is actually an improvement?
Edit: Anthropomorphizing algorithms and pattern stores doesn't really help understanding. Instead, it's apt to spread misunderstanding. Remember how long it took to purge the popular idea of "electronic brains" actually thinking, and to establish that these were restricted to executing what's actually in code? We don't need to start another level of this with "AI". (Understanding is closely related to self-awareness and consciousness, and this is dangerous ground of misunderstanding when it comes to AI. As we've seen, even staff of pioneering companies, like Google, is prone to fall for this.)
> Searle's thought experiment begins with this hypothetical premise: suppose that artificial intelligence research has succeeded in constructing a computer that behaves as if it understands Chinese. It takes Chinese characters as input and, by following the instructions of a computer program, produces other Chinese characters, which it presents as output. Suppose, says Searle, that this computer performs its task so convincingly that it comfortably passes the Turing test: it convinces a human Chinese speaker that the program is itself a live Chinese speaker. To all of the questions that the person asks, it makes appropriate responses, such that any Chinese speaker would be convinced that they are talking to another Chinese-speaking human being.
> The question Searle wants to answer is this: does the machine literally "understand" Chinese? Or is it merely simulating the ability to understand Chinese? Searle calls the first position "strong AI" and the latter "weak AI."
> Searle then supposes that he is in a closed room and has a book with an English version of the computer program, along with sufficient papers, pencils, erasers, and filing cabinets. Searle could receive Chinese characters through a slot in the door, process them according to the program's instructions, and produce Chinese characters as output, without understanding any of the content of the Chinese writing. If the computer had passed the Turing test this way, it follows, says Searle, that he would do so as well, simply by running the program manually.
> Searle asserts that there is no essential difference between the roles of the computer and himself in the experiment. Each simply follows a program, step-by-step, producing behavior that is then interpreted by the user as demonstrating intelligent conversation. However, Searle himself would not be able to understand the conversation. ("I don't speak a word of Chinese," he points out.) Therefore, he argues, it follows that the computer would not be able to understand the conversation either.
> Searle argues that, without "understanding" (or "intentionality"), we cannot describe what the machine is doing as "thinking" and, since it does not think, it does not have a "mind" in anything like the normal sense of the word. Therefore, he concludes that the "strong AI" hypothesis is false.
Even, if we don't (clearly) understand what "understanding" means, or, at least, aren't able to provide a sane definition, we do know about the semantics of the term and the kind of connotations that come with it. Like a reflexive component. (Which wasn't much of a problem in the age of behaviorism, as this had to be ignored by requirement anyway. If there is no acknowledged difference between a human and a pigeon, what is the problem with computers, as far as the model is concerned?) So we do have a notion of the semantic field and its implications. And these are, well, quite disastrous for this purpose.