Even if both aren't true, your evidence was people saying two opposing things. The truth (if there is a single objective truth on a given thing) has little bearing on whether or not different people agree on it.
I think this is due to rapidly rising expectations.
When LLMs first show they can do some new thing, we're excited at first. Then, we quickly start taking it for granted, and get upset whenever the LLM fails.
Just three years ago, LLMs could barely hold a conversation. Now, they're writing entire code bases and solving famous mathematical conjectures, but we still focus on whatever they can't do.
Something being "non-deterministic" is orthogonal to whether or not skill plays a role.
And there are benchmarks that cleanly separate the SOTA models:
Saturation of benchmarks is a property of benchmarks just as much as of the models.
- one is scaling laws, where we found years ago that pretraining validation loss scales in an almost miraculously predictable way with data volume and compute. There are apparently theoretical bases for this that I don’t quite understand but this property alone is holding at every scale we’ve ever tested. There is not just “one” scaling law but the point is there are scaling laws and they continue to faithfully predict the performance gains we see
- one is benchmarks, which I always point to epoch capability index as a good summary of them in aggregate which makes it nice to plot on one curve the capability improvement over time
To me either one without the other is substantially weaker, the fact that theory and empirical measurements give you a very good scaling law on a more unintuitive quantity (pretraining validation loss) that’s only indirectly related to the downstream performance you care about, benchmarks (in aggregate) are more direct measures of downstream performance but are harder to nail down clean and well motivated “laws” from theory (as far as I can tell). Nevertheless we do in fact see a clear trend that is not slowing.
That doesn’t mean there aren’t a whole host of benchmark problems that don’t impact the numbers involved here (leakage from training data, fundamental flaws in the design, benchmaxxing) but they don’t change the larger story. These problems don’t plausibly explain the clean trends we see.