Does seem like the contest structure could include quite a bit of risk for hiding the effect of overfitting ... I wonder if there is anything inherent about the problem that reduces that risk ...?
Does seem like the contest structure could include quite a bit of risk for hiding the effect of overfitting ... I wonder if there is anything inherent about the problem that reduces that risk ...?
The reason why the top score in one year, can be lower than in the previous year, is that the test (the 100 structures to guess) is always new and different, so it can end up being 'harder' than the year before. Luck will also play a small role.
Another explanation for a reduction in the top score would be, that previous winners are not re-submitted unchanged. For instance AlphaFold v1 seems to not have been submitted to the latest competition.
Is it really possible to select 100 new structures which together are likely to represent a meaningful increase in the sample generalization versus the prior years test set ...?
Using 1% of those (presumably from the more-often-reproduced subset) for this challenge seems reasonable? Note that the structures have to remain secret up until the challenge, and presumably all those teams uncovering the structures don't want to have to wait up to 2 years every time to actually make their results public.
I suppose it will take a few more years of repetition for the challenge to confirm that the problem has been been solved -- but I wonder if a new version of the contest is going to be needed as well? Maybe the model accuracy is now high enough to invert the contest to a form where models generate predictions for randomly selected unknown samples -- and experimental teams are then expected to make observations for those particular sequences over the next two years as part of their otherwise research agenda selected experimental workload?
But yeah, compared to other fields, the size of training/test sets is sometimes pretty small in ML for life sciences.