I heard conflicting things about it. Some claim it was trained so it can do well on benchmarks and in real world scenarios it's lacking. Can somebody deny/confirm ?
Hypothetically let's say the benchmark contains "test divisibility of this integer by n" for all n of the form 3x+1. An extremely overfit llm won't be able to code divisibility for all n not of the form 3x+1, and your benchmark will never tell.
But in modern usage it is often rephrased to: "When a measure becomes a target, it ceases to be a good measure"
https://en.m.wikipedia.org/wiki/Training,_validation,_and_te...