Background:
The technique is actually pretty fascinating. This is something that's been well understood by the cryptography community for decades, but is somehow just recently being fully appreciated by the ML community. See here:
http://blog.mrtz.org/2015/03/09/competition.html
https://www.kaggle.com/c/restaurant-revenue-prediction/forum...
Summary -
Submitting guesses to a system that gives you back scores for your guesses, will quickly leak out enough information that you can reverse engineer a huge number of hidden numbers/labels in surprisingly few iterations, e.g. 700 iterations to covertly extract 10,000+ real numbers with high precision. This surprisingly rapid convergence is a bit reminiscent of the birthday paradox.
Further, this not only lets you win the against the "test" dataset, as apposed to the final "validation" set, but this allows you to significantly increase the data available to you to train your model on, since now you can train your model against both the "test" and "training" datasets.
Layman summary -
ML breaks datasets into 3 partitions "test", "train", and "validation". In cases where they're evenly split, this technique can double the training data you have access to, which is a massive advantage in ML competitions where scores differ by tiny amounts.
Moral judgement -
My opinion, this moral argument is misdirecting the attention from where it needs to be. Yes, it's bad what occurred here. But at this point, in 2015, and with tools readily available to crack this problem effortlessly, it's inexcusable for contests to allow so many scoring reports against their validation sets anymore. It's no longer a question of whether contestants will do it, but how many of them will. We'd might as well just let people self-report their scores on an honor system, if we're going to be this overly trusting.
Try creating a contest system like this in the cryptography field any time in the past 3 decades and you'd be insulted and laughed out. Allowing so many scoring reports against the validation set is fundamentally flawed. The only solution is to globally limit calls to the scoring api.
Another proposed solution -
Allowing everyone to see everyone else's guesses & resulting scores against the "test" set, so that everyone is on equal ground for reverse engineering the "test" set, and then globally limiting the number of scoring attempts so that the test set isn't reverse engineered too significantly.
Overfitting the "validation" set actually is not a problem either way, because none of these contests are dumb enough to let anyone score against the validation set at all until the contest submission deadline is over.