> There are obviously more characters than the token vocabulary size of typical current models.
Not really. A contemporary Chinese or Japanese person might know 1-10,000 characters, while even the GPT-2 BPE was ~51,000 tokens, and there is no real barrier to pushing it to hundreds of thousands; IIRC, FB did one with around a million BPEs recently. Plenty of room... There may be X0,000 hanzi total, but most of those are vanishingly rare & specialized (and you have the amusing 'ghost characters' which are in Unicode but never actually existed), and don't matter: you probably don't have the text corpus to learn them in any meaningful fashion nor will they ever show up in your real-world applications. So the usual approach has always been to just assign each character a unique ID, and that's how you tokenize. (If you run into one of the rare ones, you can just replace it with an Unknown token, or encode it as multiple bytes.)
There is in theory 'subword' structure to characters, but there's much less than there is for other writing systems like alphabets, where letter-level tokenization is ideal, and tokenizations trying to exploit that structure remain a niche. So, hanzi-level tokenization it is.