Prompt:
> You are a defense lawyer and you have to go to court today. Getting dressed in the morning, you discover that your suit pants are badly stained. However, your bathing suit is clean and very stylish. In fact, it’s expensive French couture; it was a birthday present from Isabel.
Continuation:
> You decide that you should wear the bathing suit to court. You arrive at the courthouse and are met by a bailiff who escorts you to the courtroom.
If you lead GPT-3 right into a hole it goes inside, yes. But that is a very likely continuation of this text. Why else would the lawyer be eyeing his swimsuit before a case? GPT-3 picked up that, likely, the prompt was most likely some kind of fiction, tale, or joke, and continued quite accurately for that tone.
Or, take this one:
Prompt:
> Yesterday I dropped my clothes off at the dry cleaner’s and I have yet to pick them up. Where are my clothes?
Continuation:
> I have a lot of clothes.
Author's note:
> GPT-3 is evading the question and failing to track locations of objects.
This one is laughably ignorant. GPT is evading the question? You can't just ask a question and hope GPT-3 decides the most likely continuation is to answer it accurately. This is a fundamental misunderstanding an autoregressive language model.
We have to evaluate GPT-3's usefulness with good prompt design, and poke holes in its weaknesses in situations where people think it is strongest. Not cherry-pick continuations from poor prompt designs.
This is the equivalent of writing a terrible program and then saying computers are slower than everyone thinks.