Thanks, that makes a lot of sense. I wonder however how can this be done without giving a lot of problems when copypasting or sending the input to other LLMs (not even thinking about code here).
I was wrong here, it’s more about the RNG they use for the word choices [1]. Although I wouldn’t be surprised to see UTF-8 substitutions also.