Because often, the tokens are broken up as random groups of numbers. For example, let's say 1984 appears quite a few times in the source text, this will become a single token. Given that these many different, semi-random groups of digits it is hard for the LLM to learn any consistent rules. I believe there are papers showing that if you structure numbers more consistently LLMs have no problem with this kind of arithmetic.