That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?
This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
Isn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks
I wouldn't generalize this to all LLMs. So far I only saw Anthropic ones affected.
same here, it reads exactly the same whether the number is real or
completely made up, so the confidence stops meaning anything.