Failing at simple arithmetic after nailing some advanced physics answers has the air of playful bathos.
Failing at simple arithmetic after nailing some advanced physics answers has the air of playful bathos.
What moved me to post is that that kind of silly answer is the exact sort of shenanigans that I would pull if I were cast as the control group in a Turing test.
I already do such things winkingly when talking with my preschooler to send him epistemic tracer rounds and see if he's listening critically
that's the best phrase I've heard all year.
I do this all the time with my kids too, but I think of it more as fault injection.
AIs that don't lower the error rate are abandoned, AIs that score well are replicated and improved. It's evolution at work, but they have to enjoy (optimise for) lower error rates in order to even exist.
You of course know that the model is not capable of thought or reasoning - only the appearance of them as needed to match its training corpus. A training corpus of completely human generated data. As such, how could anything it does, be anything but anthropomorphic?
Now, if this model were trained exclusively on a corpus of mathematical proofs stripped of natural language commentary, the expectation that you seem to have would be more appropriate.
Do we know? It's the reverse Chinese room problem. :p
I aspire one day to find the free weekends and adequate hubris to build a benchtop implementation of Julian Jayne's Bicameral Mind with 1+N GPT-3 or GPT-neo instances prompting each other iteratively to see where the train of semantics wanders. (as I'm sure others have already)
I would want to ask it "15 x 7" outside of a dialogue or with examples, or look at the logprobs, or check whether "15 * 7" works (could there be something screwed up in the tokenization or data preprocessing where the 'x' breaks it? I've seen weirder artifacts from BPEs...). GPT-3 does not always 'cooperate' in prompting or dialogues or read your mind in guessing what it 'should' say, and there's no reason to expect Gopher to be any different. The space before the question mark also bothers me. Putting spaces before punctuation in Internet culture is used in a lot of unserious ways, wouldn't you agree 〜
I definitely would not hastily jump to the conclusion, based on one dialogue, "ah yes, despite its incredible performance across a wide variety of benchmarks surpassing GPT-3 by considerable margins and being expected to do better on arithmetic than GPT-3, well, I guess Gopher just can't multiply 1-digit numbers or even guess the magnitude or first digit of the result! What a pity!"
† quickly checking, GPT-3 can solve '15 x 7 =105'.
Though who knows, maybe it does have a sense of humour.
According to the link, Gopher is far better at math than GPT-3, and GPT-3 can solve "15 x 7", so I'd assume that Gopher would be able to as well.