What I mean by robust and targeted improvements is not about the concept as such, but about any choices specifically with respect to how you build the tokenization layer - if you're making a particular system, making better targeted choices for tokenization, character filter/preprocessing or vocabulary can give you some improvements in efficiency, but they rarely are a dealbraker and tokenization never is a key enabler. Like, if some tokenization or filtering destroys data your specific task happens to need, that's a problem, but you don't need advanced future research to fix it, going back to simpler tokenization and
removing features is sufficient for that, at the extreme you could always use a naive character-level tokenizer, it's trivial but simply is less computationally efficient.
If you don't care about tokenization and use any of the reasonable default options without caring about them, and if you're doing a proper pre-training on non-tiny quantities of data, then the next few layers of whatever neural architecture you have on top of these tokens will generally be able to learn to compensate for any drawbacks in your tokenization, perhaps at some computation overhead - e.g. perhaps you could have had one less layer or smaller layers if you had the best tokenization possible, and edging out that computation cost improvement is pretty much the only thing you can hope to get out of having a better tokenizer.