From some other AI demonstrations, I recall there's usually a bunch of surface-level tags with probabilities associated that are produced alongside the output. Not sure how this looks for GPT-3, but if it could provide - alongside the answer - a list of top N tokens or concepts with associated probabilities, with N set to include both those that drove the final output and those that barely fell below threshold - that would be something you could use to evaluate the result.
In the example from the article, imagine getting that original text, but also tokens-probability pairs, including: "Hobbes : 0.995", "Locke : 0.891" - and realizing that if the two names are both rated so highly and so close to each other, it might be worth it to alter the prompt[0] or do an outside-AI search to verify if the AI isn't mixing things up.
Yes, I'm advocating exposing the raw machinery to the end-users, even though it's "technical" and "complicated". IMHO, the history of all major technologies and appliances show us that people absolutely can handle the internal details, even if through magical thinking, and it's important to let them, as the prototypes of new product categories tend to have issues, bugs, and "low-hanging fruit" improvements, and users will quickly help you find all of those. Only when the problem space is sufficiently well understood it makes sense to hide the internals behind nice looking shells and abstractions.
--
EDIT: added [0].
[0] - See https://news.ycombinator.com/item?id=33869825 for an example of doing just that, and getting a better answer. This would literally be the next thing I'd try if I got the original answer and metadata similar to my example.