I think, however, we should be careful about anthropomorphizing. When the researchers wrote 'inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”', did they have evidence that this was being attempted, or are they thinking that if a person made this error, it could be explained by their not carrying a 1?
I also think a more thorough search of the training data is desirable, given that if GPT-3 had somehow figured out any sort of rule for arithmetic (even if erroneous) it would be a big deal, IMHO. To start with, what about 'NUM1 and NUM2 equals NUM3'? I would think any occurrence of NUM1, NUM2 and NUM3 (for both the right and wrong answers) in close proximity would warrant investigation.
Also, while I have no issue with the claim that 'the best way to accurately predict an arithmetic answer is to deduce how the mathematical rules work', it is not evidence that this actually happened: after all, the best way for a lion to catch a zebra would be an automatic rifle. We would at least want to consider whether this is within the capabilities of the methods used in GPT-3, before we make arguments for it probably being what happened.