Let's test GPT-2: https://huggingface.co/gpt2?text=One+apple+is+1%2C+two+apple...
Prompt: "One apple is 1, two apples are 2, three apples are" Model output: "One apple is 1, two apples are 2, three apples are 4, four apples are 5, six apples are 7, seven apples are 8, nine apples are 10, ten apples are 11 (for apples being the perfect length of life)."
Even if you use a special dataset to each it to count, it won't be able to count beyound the examples in the training set. So it's spewing plausible-sounding gibberish at you (i.e. approximating the training set distribution)
It doesn't generalize. Not in the sense that it can't give you a phrase that didn't exist in the training set. It can. But it can't give you a new kind of phrase, of a "kind" that didn't exist in the training set.