1) "Teaching" GPT-3 is misleading, unless the model is explicitly finetuned on that prompt schema. The default davinci model for the OpenAI API is the same model for everyone using it; the few-shot prompts merely guide GPT-3 toward a domain and generation schema.
2) This approach is essentially a classification problem, with extra steps. That's not new with GPT-3; causal language models (including GPT-2) use classification tasks as a way to quantify model performance (see the papers). Hugging Face transformers is a good library to finetune Transformers specifically for classification tasks.
3) The full string "yo be real" is inefficient if you're just checking for positive/negative responses. (Token efficiency will likely be an interesting area of optimizing GPT-3 performance)
4) That said, using token prediction probabilities can be useful for assessing model performance (I explicitly duplicated the web UI's implementation for my new client: https://twitter.com/minimaxir/status/1288125135526875146). However, the prediction probability is often very dependent on previous tokens, including the prompt, so the results are not universally consistent. But there might be more clever outside-the-box approaches using the token probabilities and heuristics...
tbh, I wish OpenAI would also return attention weights for each previous token in order to be able to more easily-visualize which tokens have a high influence on generation.