I wouldn't disbelieve that the grouped version is actually better with data, but it fights my intuition pretty hard. Grouping based on frequency obfuscates the regular nature of numbers.
I wouldn't disbelieve that the grouped version is actually better with data, but it fights my intuition pretty hard. Grouping based on frequency obfuscates the regular nature of numbers.
Yeah
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs - https://arxiv.org/abs/2402.14903
xVal: A Continuous Number Encoding for Large Language Models - https://arxiv.org/abs/2310.02989
I believe there's another paper that demonstrates something like also for the likes of spelling, counting etc but i can't remember it.
Very interesting paper. It does make sense to me the R2L chunking would be better than L2R chunking. It doesn't actually study single digit tokenization.
I am mostly interested in a direct comparison between an LLM wide tokenization vs single digit tokenization. It would be nice to see a direct comparison between similarly trained models. Otherwise it is very hard to get a definitive answer by comparing models with varying sizes, training time, and general strength.
> xVal: A Continuous Number Encoding for Large Language Models - https://arxiv.org/abs/2310.02989
I have seen this paper before, but hadn't payed attention to the p10 vs p100 analysis. Its not clear that the findings would be relevant to an LLM like gtp4 though.
It's literally a lookup table to go from token to embedding. What would you expect as improved? At best, maybe cache coherency if groups of numbers are converted to embeddings sequentially... But embeddings are huge (e.g. 8kb for llama-2) so you're losing caches jumping around between non-contiguous numbers anyways.
For example in order for a super simple model to learn 3 digit multiplication, it would need to see at least one example for each token in order to get ANY information about what number it represents. Alternatively, with single digits you only need an example where each position is present in each location. Obviously, we would hope to have plenty of data, but I would expect better generalization from models which need to rely on memorization less.
Alternatively, I can see a few reason why grouped digits would be better, but they are more complicated reasons than the reason above so by Occam's Razor my intuition says single digits should be better.