Try searching for different words using the search box here: https://observablehq.com/@simonw/gpt-tokenizer#cell-135
This could force the model to correctly learn how to capitalise, make all-caps, etc…
The goal is simply to speed up training slightly, it wouldn't actually make a difference to the final performance of a model as big as GPT-4 (except maybe decrease the prevalence of glitch tokens)
Doesn't that assume that the embeddings learned are in some sense "perfect"? Is that actually the case in practice?
I would expect the learned embeddings to have some errors, especially for the rarer ones that have few examples available for the model to learn from.
I also thought that explicitly accounting for symmetries always improved model performance, because then it doesn't waste parameters learning things that aren't unique and interesting pieces of information.