How would we train it? Don't we need it to understand the heaps and heaps of data we already have "tokenized" e.g. the internet? Written words for humans? Genuinely curious how we could approach it differently?
OpenAI's tokenizer makes "chess" "ch" and "ess". We could just make it into "c" "h" "e" "s" "s"
That is, the groups are encoding something the model doesn't have to learn.
This is not much astray from "sight words" we teach kids.
Yup. Just let the actual ML git gud
Worth noting that the relationship between characters to token ratio is probably quadratic or cubic or some other polynomial. So the difference in terms of computational difficulty is probably huge when compared to a character per token.
There is no advantage to tokenization, it just helps solve limitations in context windows and training.