I read recently that small variations in the tests cause failures by large margins.
If this doesn’t show over fitting in don’t know what would.
If this doesn’t show over fitting in don’t know what would.
The math one in particular is the one where small variations reduce the success rate significantly. I can’t find the source but it was pasted here in the last 2 weeks.