It's similar to how it is splitting words at arbitrary points, rather than at clear morphological or lexical locations (e.g. on the Jane Austen text `"Now, ma'am," said Jane to her aunt, "shall we join Mrs. Elton?"` I've seen it tokenize that as `"|Now|,| ma|'|am|,"| said| Jane| to| her| aunt|,| "|shall| we| join| Mrs|.| El|ton|?"`).
Moreover it's trivially easy to tokenize the glyphs.
BPE is a tradeoff between single letters (computationally hard) and a word dictionary (can't handle novel words, languages or complex structures like code syntax). Note that tokens must be hardcoded because the neural network has an output layer consisting of neurons one-to-one mapped to the tokens (and the predicted word is the most activated neuron).
Human brains roughly do the same thing - that's why we have syllables as a tradeoff between letters and words.
Yes, I guess the point here is that the glyph, not the byte, is the base unit of communication in Unicode charsets.