GPT-4 (I haven't really tested other models) is surprisingly adept at "learning" from examples provided as part of the prompt. This could be due to the same underlying mechanism.
It might be more effective to try to play 'tokendle' before trying to play 'wordle'.
Or would an LLM get confused if we were to alter the way the tokenization of the input text is done, since it probably never encountered other token-"spellings" of the same word?
https://www.geeksforgeeks.org/lzw-lempel-ziv-welch-compressi...
For 'code table' substitute 'token table'.
It's basically unrelated to what happens during training, which is using gradients.