Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.
Byte Latent Transformer - https://arxiv.org/abs/2412.09871
1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.
Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.
I've got no problem with an AI doing something similar
If it was top priority, every company that can't find a post training fix would go disable half their tokenizer code and it would be solved in the next model.