> So, in general Claude still beats newest GPT-4o model in software development category, right? [1]
I'm genuinely curious to which extent widely published benchmarks can be used for this kind of assessment. The exercises this benchmark uses seem to be available in https://github.com/exercism/python/tree/main/exercises, and it seems that there are solutions at least at https://exercism.org/tracks/python/exercises/series/solution....
Overfitting used to be a concern when it came to comparing ML models, but it feels like it's not being talked about that much anymore, is it? What if one model has the exercism solutions in its training data, but the other model doesn't, is the comparison still fair? If it solves 70% of the exercism exercises correctly, would it perform as well on similar prompting for unseen problems (which are more likely to be what I care about)?