"nobody is disabling this optimization" shouldn't be surprising. It's a matter of knowing what letters are in words versus a 4x token efficiency boost. Nobody cares enough about spelling.
If it was top priority, every company that can't find a post training fix would go disable half their tokenizer code and it would be solved in the next model.