I haven't read the paper so I'm not sure if they did this, but it would be interesting to see at what point it breaks down. Just scrambling up letters within words makes it pretty easy for the LLM; what if you also start moving letters between words, or take out the spaces between words?
Or perhaps they inserted typos automatically in the training set as data augmentation. Tactics like that is known to increase the roboustness of some models, so why not?
It is trained on data which may include typos, but that is very different from fixing typos. It knows what words likely come after typos in the same way it knows what words likely come after regular words.
Even non-finetuned 7B models, 3 orders of magnitude smaller than GPT-4, can unscramble text and fix typos reliably.
Half, or better, of the things people discover "GPT-4 can do" can be done with non-RLHF GPT-3 from 2020 or with a model 1000x smaller.