I have just a question: how many prompts did it completely fail for each prompt it answered brilliantly? 5? 10? 100? That's what these "twitter threads showing gpt3 doing astonishing things" fail to show.
The initial problem that got me interested in GPT-3 is suppressing confabulated answers to the Hofstadter-Bender questions published in The Economist. I eventually found an apparent solution but I’m yet to carefully validate it: https://twitter.com/goodside/status/1556459121834168320?s=21...
The hallucination-suppressing prompt it implicitly uses the output of is here: https://twitter.com/goodside/status/1556459121834168320?s=21...