sure, you can get the model to understand with enough examples, but OP of this thread was saying LLMs need many many more steps than humans do. Mschuster was saying this could partially be an artifact of the tokenization - humans visually seeing 1,000,000 and 1000001 for the first time can see that they are similar based on how they look. A transformer must learn two token sequences, which may not have any tokens in common, are different by 1.* And it must do this based entirely on context and how likely one is to appear in the same place as another. That takes a massive amount of data in practice and could be one reason why humans appear to learn from much sparser data and faster.
* in gpt-2’s case, you get
1,000,000 -> 16, 11, 830, 11, 830, 198, 49
And
1000001 ->
388, 486
which means the model will see a sequence of noise embeddings entirely different when training. That means seeing 1,000,000 any number of times will not really help the model prepare for when it sees 1000001 for the first time.
You can play around with https://tiktokenizer.vercel.app/?model=gpt2 to see how tokenization of numbers has been improved by different models to try and mitigate this.