Part of the problem here is that GPT-3 has such a small vocabulary. It's 50K tokens, and many of those are either garbage, punctuation, or full words (rather than sub words).
I'd be curious to see what scaling up the size of the vocabulary would do to improve these results in a model like GPT-3...