I don't appreciate your snark at all. I made a mistake and didn't read the
paper carefully again so I confused myself with what they mean by one- two-
and three-digit tasks, I accept that. I don't see what you get from pouncing
on my mistake, other than a few marks for an internet burn.
Now, the two- and three digit addition and subtraction tasks (operations on
numbers between 0 and 99 and 0 and 999, respectively) are both small enough
for the large, 175B parameter model to have memorised them exactly. Even if
there was a single parameter for each three-digit number, of which there are a
million, you could fit the entire set 175 thousand times in the 175 billion
model (assuming they mean "a billion" as "one thousand million", not "one
million million", which they don't clarify, but to be on the safe side let's
assume the smallest). There is plenty of room.
These four tasks are also the tasks that are most likely to be present in
their entirety in a corpus of natural language, as the one GPT-3 was trained
on, for example as records of common monetary transactions (especially the
two-digit ones). That is, yes, the training set can comfortably contain 800k
3-digit addition problems. Why not? It contained 410 billion tokens from the
Common Crawl dataset alone, plus a few extras.
In short, the almost perfect accuracy on this task is not impressive. The 25%
ish accuracy on the four-digit addition task is even less impressive. I don't
know what the baseline is here, but 25% accuracy on anything is not something
to write home about.
You ask me to provide evidence of my own to support the memorisation claim.
The claim is not memorisation. The claim is that GPT-3 has learned arithmetic
(not stated exacly like that in the paper). This claim flies in the face of
the commonly understood operation of language models, which are systems that
compute the probability of a token to follow a sequnce of tokens- and nothing
else. It's very hard to see how such a system should be able to perform
arithmetic operations, while it's very easy to see how it can instead memorise
their results. If the authors of the GPT-3 paper wish to claim that GPT-3 can
perform arithmetic, instead of the much simpler explanation, they have to
provide very strong evidence to back that up and refute the simpler
explanation.
And the "spot checks" that they performed are nowhere near such strong
evidence: I can fail to find anything I search for, if I search with the wrong
terms and the authors don't give much information about how they did their
"spot checks". I mean, did they use a regular expression? Which one? ("<NUM1>
+ <NUM2> =" is not a regular expression! But then - what is it?) Did they take
into account whitespace? Punctuation? Something else? What search terms they
used? They dont' say. Can we tell why they failed to find what they were
looking for? No.
Besides, why only "spot check" three-digit arithmetic? It would make a lot
more sense to spot-check two-digit problems, first, because these are the
most likely to be found more often in the dataset and consequently be
memorised. Indeed, the fact that they don't report "spot checks" for two-digit
arithmetic suggests that they did perform those spot checks and they found a
lot more overlap than for the three digit arithmetic, but chose not to report
it. And if their model was memorising two-digit arithmetic, and that explains
its performance on that type of task, it's safe to assume that it was
memorising the third-digit arithmetic task also and that their "spot checks"
were simply not very well put together to find the three-digit arithmetic
examples.
Note that section 4 goes in length over the possibility that the test set for
all tasks (not just arithmetic) was contaminated (i.e. that it containted
training examples from existing benchmarks, published on the internet). I
haven't read that one carefully but test set contamination is another
possibility. And, to be frank, any possibility is more possible than the
possibility that a langauge model has learned arithmetic- which is tantamount
to magick.