> When using GPT-2 vs GPT-3 vs GPT-4, a human can easily tell that each is leaps and bounds "better" than its predecessor, with "deeper" understanding of the input and "more human-like" responses and reasoning. There is a strong impression that those models aren't just progressing along a scale, but changing qualitatively.
I agree with you about all of this.
> I simply doubt that any of the proposed metrics captures this intuitively observable quality. Furthermore, I claim that it is this quality that actually matters.
I think there are important qualitative observations about LLMs that create understanding and guide future research, but I think it's bad science/engineering to rely on some indescribable quality of "goodness" when it comes to developing new models.
Metrics may not be perfect, but they are all we have. You could train a model, send it to a bunch of human crowdworkers, and ask them to rate it on how good it is (Anthropic does this!), but the result of that is... another metric.
That said, yeah, I think we definitely could use better metrics. OpenAI agrees, which is why they're pushing their Evals project super hard. The paper we're commenting on agrees, which is why they're proposing alternatives to these step-function metrics.
> We don't need a language model in order to multiply two numbers. Any clearly defined and algorithmically solvable (and thus readily quantifiable) task is trivial for regular software.
A language model multiplying two numbers is the whole point of emergence.
Transformers were designed to translate between different natural languages and trained on completing sentences with some words masked out. As we scaled them bigger and bigger, we found that the same model architecture was suddenly capable of more than completing sentences: it was answering questions about high school geography (MMLU), writing computer programs, and - yes - doing basic arithmetic. New capabilities were emerging with scale.
This means that a "language model," at sufficient scale, is more than a language model. This is the closest thing we have right now to generalized AI, where one model can complete a variety of disparate tasks. The arithmetic thing is exciting and unexpected because of what it represents, and understanding it should be a priority, even if we have better non-ML ways of doing that particular task.