Some (all ? ) Kaggle competition also have a daily submission limit to avoid this kind of cheating.
People can, and do, overfit their model to the public test set, but doing so does not improve their score on the private test set, so cheating is prevented even without the submission limit.
The submission limit helps ensure the leaderboard generated from the public test set stays close to the leaderboard generated by the private test set while the competition is running, so that you can get an idea about your standing. But participants know better than taking the public leaderboard too seriously.
This is what the Baidu team were probing.
Alternatively, require the participants to provide an API and call them, instead of them submitting things.
But if you can repeatedly run different models against the test set and you do get to see a score, you could do something like random parameter searching or other optimization ideas to tune your algorithm to be highly overfitted to the test set.
Another way to avoid this would be to develop several test sets that are roughly "equivalent" in terms of the distributional properties, and then randomly change the test set periodically, or change the test set right after the final submission deadline, to discourage people from pursuing overfitting.