Neat. Is it a single under-trained token in GPT-5.2? Or is something else going on?
https://raw.githubusercontent.com/niieani/gpt-tokenizer/refs...
Indeed, how do they deal with Chinese? Are some ideograms multiple tokens?
(ges)(chn)(iegelt)
[1]: https://platform.openai.com/tokenizer"This word is geschniegelt" is [2500, 2195, 382, 192786]
Last token here is " geschniegelt"