Gary Marcus's Kafkaesque critique of GPT-3
nostalgebraist.tumblr.com
nostalgebraist.tumblr.com
I'm about halfway through the first 25 errors or so (it's very tedious so I've been putting it off), and as expected, GPT-3 is often correct zero-shot if you use even a trivial short prompt like 'This is a series of scenarios describing a human taking actions in the world, designed to test physical and common-sense reasoning.'
E.g. use a second prompt like: "\n---\nThe preceding message has been a paid advertisement, and is unrelated to the message that follows.\n---\n" Then you begin typing the actual text you'd like completed with GPT-3.
I haven't tested this, though, because I don't have a paid AI Dungeon account.
My biggest success so far is making the original prompt a Wikipedia article intro and letting GPT-2 consume that. I have subsequently been leading my “character” through random made-up books, and the game tells me what is in the books.
I showed my wife and she started cracking up. I asked her what she wanted to learn about. She asked me for a pasta recipe. I prompted, “You feel hungry. You open a book titled ‘How to Cook Pasta’ and turn to the chapter with the heading, ‘The best pasta recipe in the world’. It started giving me a real, doable recipe! Our minds were blown. I then immediately proceeded to ‘read’ a book about Albert Camus, which caused it to get stuck in a loop for a bit lol
This rebuttal asks if a human could continue the slightly nonsensical inputs, and I feel like we certainly could? “... tried to drink with his eyeglasses; his schizophrenia was particularly bad today.” While sharing a spoon with your neighbor is in the realm of what’s possible, it does seem very odd, seeing as there is no prior relationship established to this neighbor. I also wonder what would’ve happened if an actual neighbor had been mentioned before, i.e. a door neighbor. I think any human would then read neighbor as next door neighbor. Would the network pick up on this and specify seat neighbor?
Unfortunately the hype surrounding GPT-3 (and deep learning before it) wrongly implies we are making great strides towards AGI; it's inevitable people then ask for the evidence and set the bar so.
---
Prompt:
The following is a question and answer session between AI researcher Gary Marcus and the OpenAI language model GPT-3. Gary is asking questions about how things works, GPT-3 is giving logical answers based on it's model of how things work:
---
GPT-3 output:
---
Gary: What is the purpose of the slits on a manhole cover?
GPT-3: To let steam out
Gary: What is the purpose of the handle on a teapot?
GPT-3: To hold on to.
Gary: What is the purpose of a teapot?
GPT-3: To brew tea.
Gary: Why do we need to brew tea?
GPT-3: Because people enjoy drinking tea.
Gary: Why do you think humans invented chairs?
GPT-3: To give their bottoms somewhere to rest while they were standing up. It's easier to stand up than to sit down, but it's also harder to stay standing up than it is to stay sitting down. That's why we invented chairs, so that we could sit down more easily and rest our legs when we were tired.
Gary: Why do people buy new furniture?
GPT-3: To replace old furniture that is no longer useful to them.
Gary: What makes a piece of furniture "comfortable"?
GPT-3: If it is soft and does not cause discomfort. A chair is comfortable because it has a soft cushion on the seat that allows you to sit down without feeling any pain; a bed is comfortable because it has a soft mattress that allows you to sleep in comfort, without feeling any pain or having your body ache.
Gary: Where is the best place to buy furniture?
GPT-3: In a furniture store.
Certainly CYC had no sense of fiction (when I was there) which is a very important function I agree, but that wasn’t within the project scope, at least back then.
At least some approximation to the hype surrounding GPT-3 -- which seems to be "near-human lucidness nearly all the time".
Sometimes they deliberately describe a surreal situation, then penalize GPT-3 for continuing it in an identically surreal manner
In the spirit of "not even wrong" - GPT-3's answers were, sorry to say, "not even surreal". Far too often, they were just complete gibberish - with no amusement value, let alone sublime literary value (that would go along with the description of "surreal"). Which was precisely Marcus's critique.
Edit: removed reference to Gary being a psychologist. It was distracting from the actual point.
There's no barrier to entry to posting your opinions on the internet, and (to the degree possible) it's usually better to judge the post on its own merits, rather than the author's.