This isn’t actually true, and is a persistent myth. Or rather, you should back up the claims with evidence.
It’s a bit like saying that you perform poorly on character manipulation tasks because you don’t read individual letters.
Biology analogies aside, I haven’t seen anything to suggest that utf8 level tokenization causes a significant decrease in perplexity across large datasets. (Note that the “large dataset” criteria is required. It’s certainly possible to demonstrate improvements in restricted cases, but no one is really interested in the restricted case unless you have a very specialized task. In which case, sure, specializations make sense — ChessGPT being an obvious example where tokenization just harms learning.)
So the tradeoff isn’t the large context window, but rather the desire to have a deep understanding of a massive amount of data. Specialized models will always have a place as a small component of the whole, but suggesting that this is a problem solved by superior architectures seems a little bit of a stretch.
I think what’s going on here is that OpenAI spent a lot of time giving feedback to their model about specific use cases, and ROT-13 was obscure enough (both in usage and in the data) that its performance is limited. I’d bet that if OpenAI did a few rounds of RL on this objective, the model would perform as well as its cousins.