It's not "some new kind of architecture comes along". It's been done before many many times and it's a very basic change, turning off a specific optimization. And it's a perfectly good style of LLM so "inherent flaw" isn't true either.
If it was top priority, every company that can't find a post training fix would go disable half their tokenizer code and it would be solved in the next model.