> How many "x + y" questions can be formulated where x and y are both single-digit numbers? The answer is 2^10, or 100.
Less snarkily, if there’s (10^4)^2 = 100 million combinations of 4 digit addition problems, and GPT-3 is reaching 25.5% accuracy on those problems (vs. 0.4% in the 13B parameter model). For 3 digit problems, it’s even better: 1 million combinations and 80.4% accuracy. Clearly, there is more happening than simple memorization—the training set does not contain 800k 3-digit addition problems. Thus, it’s fair to say that the model has at least a partial grasp of how to perform arithmetic operations (but probably not fair to say that it has synthesized the entire system of arithmetic).
Also, the paper does say that they scrubbed exact examples from the training set to avoid memorization, a fact you left out:
> (pg. 23): ”To spot-check whether the model is simply memorizing specific arithmetic problems, we took the 3-digit arithmetic problems in our test set and searched for them in our training data in both the forms "<NUM1> + <NUM2> =" and "<NUM1> plus <NUM2>". Out of 2,000 addition problems we found only 17 matches (0.8%) and out of 2,000 subtraction problems we found only 2 matches (0.1%), suggesting that only a trivial fraction of the correct answers could have been memorized. In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather than memorizing a table.”
This is super interesting and something I hadn't read before. That is very cool, and definitely suggests it's figuring out how the computation actually works (!).
The nature of the failure (errors where it 'forgot' to carry the one) suggest that it's doing something like basic arithmetic and making mistakes.
This is evidence in the direction of having a model of how to do basic arithmetic and evidence against memorization.
I'm not pretending that both outcomes mean it knows arithmetic. For example, if the outputs were random or if they only matched exact examples it had seen then it would look like memorization, but that isn't what's seen.
In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather thanmemorizing a table.
So, what is "often"? 100% of the time? 60% of the time? 30% of the time? Such a vague statement is no evidence of anything, much less the very strong claim made in the paper.
Now, the two- and three digit addition and subtraction tasks (operations on numbers between 0 and 99 and 0 and 999, respectively) are both small enough for the large, 175B parameter model to have memorised them exactly. Even if there was a single parameter for each three-digit number, of which there are a million, you could fit the entire set 175 thousand times in the 175 billion model (assuming they mean "a billion" as "one thousand million", not "one million million", which they don't clarify, but to be on the safe side let's assume the smallest). There is plenty of room.
These four tasks are also the tasks that are most likely to be present in their entirety in a corpus of natural language, as the one GPT-3 was trained on, for example as records of common monetary transactions (especially the two-digit ones). That is, yes, the training set can comfortably contain 800k 3-digit addition problems. Why not? It contained 410 billion tokens from the Common Crawl dataset alone, plus a few extras.
In short, the almost perfect accuracy on this task is not impressive. The 25% ish accuracy on the four-digit addition task is even less impressive. I don't know what the baseline is here, but 25% accuracy on anything is not something to write home about.
You ask me to provide evidence of my own to support the memorisation claim. The claim is not memorisation. The claim is that GPT-3 has learned arithmetic (not stated exacly like that in the paper). This claim flies in the face of the commonly understood operation of language models, which are systems that compute the probability of a token to follow a sequnce of tokens- and nothing else. It's very hard to see how such a system should be able to perform arithmetic operations, while it's very easy to see how it can instead memorise their results. If the authors of the GPT-3 paper wish to claim that GPT-3 can perform arithmetic, instead of the much simpler explanation, they have to provide very strong evidence to back that up and refute the simpler explanation.
And the "spot checks" that they performed are nowhere near such strong evidence: I can fail to find anything I search for, if I search with the wrong terms and the authors don't give much information about how they did their "spot checks". I mean, did they use a regular expression? Which one? ("<NUM1> + <NUM2> =" is not a regular expression! But then - what is it?) Did they take into account whitespace? Punctuation? Something else? What search terms they used? They dont' say. Can we tell why they failed to find what they were looking for? No.
Besides, why only "spot check" three-digit arithmetic? It would make a lot more sense to spot-check two-digit problems, first, because these are the most likely to be found more often in the dataset and consequently be memorised. Indeed, the fact that they don't report "spot checks" for two-digit arithmetic suggests that they did perform those spot checks and they found a lot more overlap than for the three digit arithmetic, but chose not to report it. And if their model was memorising two-digit arithmetic, and that explains its performance on that type of task, it's safe to assume that it was memorising the third-digit arithmetic task also and that their "spot checks" were simply not very well put together to find the three-digit arithmetic examples.
Note that section 4 goes in length over the possibility that the test set for all tasks (not just arithmetic) was contaminated (i.e. that it containted training examples from existing benchmarks, published on the internet). I haven't read that one carefully but test set contamination is another possibility. And, to be frank, any possibility is more possible than the possibility that a langauge model has learned arithmetic- which is tantamount to magick.
The claim that “GPT-3’s performance on arithmetic tasks is solely due to memorization / data leakage—it has no generalization ability on this type of task”, is easily attackable by...well, doing arithmetic and applying common sense.
There are 2(10^5)^2 = 20 billion possible 5 digit problems (both addition and subtraction). The accuracy on those tasks is about 10%, so roughly 2 billion* 5-digit addition and subtraction problems would need to be represented in the training data (Common Crawl + books as you said). Each problem is at minimum 5 tokens (e.g. 99999 + 11111 = 111110). So is ~2.5% of the training corpus 5-digit addition and subtraction problems that eluded their filtration process? (assuming it’s ~400B tokens like you said). Seems exceedingly unlikely, so much so that memorization ceases to be the simplest explanation.
That said, yes it is surprising that a language model can generalize in this way—that’s the point of the paper. How exactly this happens seems like a valuable thread to pull. Your critiques may help, but writing the results off as impossible magic does not.
About the 5-digit problems- I didn't make this clear but I don't think those were memorised. I think the two- and three-digit problems (all three operations) were memorised, because those are the most likely to be represented in their entirety, or close, in GPT-3's training corpus, given that they are operations that are common to very common in daily life.
I doubt that the four- and five-digit addition problems (and the single-digit, multi-op problem) were represented often enough in GPT-3's training corpus for them to be memorised. I think the low accuracy in these problems (less than 10% in the few-shot setting and near zero in the zero- and one-shot) is low enough that it doesn't require an explanation other than a mix of luck and overfitting that is common enough in machine learning algorithms that it's no surprise. e.g. we evaluate classifiers using diverse metrics, not just accuracy, because this is so common.
It is this observation, that GPT-3 did well in problems that are likely to be well reprsented in its training corpus and badly in ones that aren't, that convinces me that no more complicated explanation is needed than memorisation.
Something else. Like I say above, we evaluate classifiers not only by accuracy (the rate of correct answers), because accuracy can be misleading. e.g. a classifier can have 100% accuracy with 0% false positives and 100% false negatives. The GPT-3 authors only tested the ability of their model to give answers to problems stated as "x + y = ". They didn't test, e.g. what happens if they prompt it with "10 + 20 = 40, 38 + 25 = ". Testing for aberrant answers following from such confusing prompts has often showed that language models that appear to be answering questions correctly because of a deep understanding of language are in truth overfitting to surface statistical regularities. See for example [1,2] and many other references in [3].
Indeed, I could be wrong about rote memorisation and GPT-3 can still not be learning to perform arithmetic computations, given the tendency of language models to learn spurious correlations. There is an article about a mathemagician on the front page today, that shows how she found roots of huge numbers by finding shortcuts around expensive calculations. For instance, all sums between numbers ending in 5 end in 0, etc. I wouldn't find it magickal if a language model was finding such heuristics and that this is the "something else" that is said to be going on. However that would not be "generalisation" and it would not be learning to perform arithmetic.
In the end, I don't understand how a model can be said to know how to add two-digit numbers perfectly but not five-digit numbers. If it's performing an incomplete computation in the latter case, then what kind of incomplete computation is it performing? If it "gets it wrong after three digits" then why does it get three digits right? What's the big difference between three- and four-digit numbers that causes performance to fall off a cliff - other than the chance of finding such numbers in a natural language corpus?
As to magick- I'm writing off not the results, but the hand-waving presented in place of an explanation as magick. GPT-3 is a technological artifact designed to do one job, now reported to be doing another. This requires a thorough explanation but instead we got magickal thinking: the authors wish that their models could learn arithmetic, so they took its behaviour as proof that it learned arithmetic.
___________________
[1] Probing Neural Network Comprehension of Natural Language Arguments
https://www.aclweb.org/anthology/P19-1459/
[2] Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
https://www.aclweb.org/anthology/P19-1334/
[3] https://www.technologyreview.com/2020/07/31/1005876/natural-... (try F9 if you 're over limit)
My point was that intelligent humans can and often do make mistakes in logic and computation (arithmetic) in ways that machines typically do not. One reason may be colliding or incomplete representations of certain concepts, and (relatedly) the fact that we are relying on language. I think of neural networks as fuzzy representation composers, so it seems they also fail for similar reasons. Basically, it (GPT) does have some layered representation of the concept of numbers and how they are used in different contexts which gives it some faculty at carrying out common operations, but it doesn’t “add up” to a reliable system of logic (that would allow it to extend addition to say 100-digit numbers, the way even a sharp and/or patient 2nd grader could do, generalizing from the simpler cases).
I think accuracy is sometimes the correct measure, and in this instance it seems fine—at baseline, we should expect ~0% accuracy since it is generating output from essentially the space of all possible text (texts <= 2048 tokens). I agree that it would be interesting to probe the model with better tests, and understanding when/why it fails on certain arithmetic problems or types of reasoning.
I liked what you wrote about finding heuristics, though I disagree with your conclusion that heuristic finding does not qualify as learning—it is just somewhere along the spectrum between a randomized model and an ALU (neither of which can be said to have learned anything) in terms of its ability to perform arithmetic.
Of course, we already have better models for solving proofs and such, so I generally think the way toward more complete AI models is to return to the system design view of AI (meta-learning, integration of different models, etc) rather than trying to evolve one colossal model to rule them all. That is, a meta-model that recognizes what sort of problem it is facing, then selecting a model/program to solve or generate possible solutions to that problem, while revealing or explaining as much of this process as possible to the user.
In any case, I have definitely have more to read on the subject and am mostly musing at this point. Thanks for the references and the conversation.
I guess I can concede that the memorisation explanation is not the only possible one, there's always the possibility of learned heuristics. I still expect very strong evidence before I'm convinced that GPT-3 can learn arithmetic in the general sense and I don't trust the explanation that it's only learning partially- but let's agree to disagree on that. Thank you for the conversation, too.