I was surprised to hear Dr. Bubeck, who is in kind of a privileged position wrt OpenAI and obviously an accomplished scientist, essentially saying things like "I tried asking it X and I am pretty sure it's not in the training set, and it worked, therefore I think it understands".
A really big problem with the anecdata approach to proving out AI is alluded to in Dr. Bender's story about the "Everything in the Whole Wide World" museum (for those who didn't watch, Grover the muppet goes to the aforementioned museum, sees many things — but not "everything", then walks through a door labeled "Everything Else" that leads outside). That is, no individual prompt-and-response can be relied on to inform about another prompt-and-response that you haven't yet tried. And with nondeterminism, you can't even rely on that prompt-and-response to remain stable. But a failed prompt also carries the possibility that a tweaked prompt could produce a success. So (as we saw with the twitter post yesterday complaining about LLMs not working well that was both heavily upvoted and heavily contested in the comments), we are in a world where nobody can say much except to provide anecdotes of the thing failing or succeeding or having characteristic Y, which are exercises in narrative construction to support or oppose investment, not principled discussion.