How the BPE tokenization algorithm used by large language models works
sidsite.com
sidsite.com
So, how do models like GPT answer in Chinese? Are they able to produce any Chinese character? From what I understand, they are not.
My second question would then be, which tokenization algorithms are used for Chinese and other East Asian languages. What does that mean for the models? How do models that can learn proper Chinese (with complete tokenization) differ from the models for languages with less characters?
Not really. A contemporary Chinese or Japanese person might know 1-10,000 characters, while even the GPT-2 BPE was ~51,000 tokens, and there is no real barrier to pushing it to hundreds of thousands; IIRC, FB did one with around a million BPEs recently. Plenty of room... There may be X0,000 hanzi total, but most of those are vanishingly rare & specialized (and you have the amusing 'ghost characters' which are in Unicode but never actually existed), and don't matter: you probably don't have the text corpus to learn them in any meaningful fashion nor will they ever show up in your real-world applications. So the usual approach has always been to just assign each character a unique ID, and that's how you tokenize. (If you run into one of the rare ones, you can just replace it with an Unknown token, or encode it as multiple bytes.)
There is in theory 'subword' structure to characters, but there's much less than there is for other writing systems like alphabets, where letter-level tokenization is ideal, and tokenizations trying to exploit that structure remain a niche. So, hanzi-level tokenization it is.
I'm pretty sure the set of tokens also contains all 256 Bytes to cover such cases.
Unicode 15 has nearly 150000 characters and CJK languages have even more than that because of Han unification.
A model like GPT-3 can only output a very primitive version of Chinese. My question is how real Chinese models deal with this and specifically how tokenization works in that case.
And also suddenly the B in BPE makes a lot of sense.