"I took the water bottle out of the backpack so that it would be [lighter/handy]"
What is lighter and what is handy? No amount of stochastic language manipulation gets you the answer, you need to understand some rudimentary physics to answer the question, and as a precondition, you need a grammar or ontology.
It sounds like you're saying "It doesn't work because it can't work", but you haven't actually shown that it doesn't work.
In the example above it guesses wrongly, but again this is not surprising because it can't possibly get the right answer (other than by chance). The solution here cannot be found by correlating syntax, you can only answer the question if you understand the meaning of the sentence. That's what these schemas are constructed for.
edit: Retracted a test where it seemed to know which to select, because further tries revealed it was random.
edit: I did some more tries, and it does seem to be somewhat random, but the way it continues the sentence does seem to indicate that it has some form of operational model. It's just hard to prompt it in a way that it is "forced" to reveal which of the two it's talking about. Also, it seems to me its coherence range is too short in GPT-2. I would love to try this with GPT-3.
The paper evaluated Winograds: https://arxiv.org/pdf/2005.14165.pdf#page=16
But with the speed of the field, maybe we can figure it out in three years. It just seems like we're still missing some key components. Primarily, reasoning and learning causality.
many people have different definitions for AGI though. for me it clicked when i realized that text has this universality property of capturing any intent.
If you really probe at GPT, you'll see anything that goes beyond an initial sentence or two really starts to show how it's purely superficial in terms of understanding & intelligence; it's basically a really amazing version of Searle's Chinese room argument.
I also would contend that there is reasoning happening and that zero-shot demonstrates this. Specifically, reasoning about the intent of the prompt. The fact that you get this simply by building a general-purpose text model is a surprise to me.
Something I haven't seen yet is a model simulate the mind of the questioner, the way humans do, over time (minutes, days, years).
In 3 years, I'll ping you :) Already made a calendar reminder
I look forward to this ping :)
From what I understand, its not just that the GPT-3 has impressive performance but more what is signifies and that is the fact that massive scaling didn't produce diminishing return, and if this pattern persists, it can get them to the finish line.