The public leaderboard gives some feedback on your model performance. But when a human is in the feedback loop, then there is a risk of overfitting. Overfitting can be explained basically as: "memorizing the data" or "learning from noise, not signal".
When you overfit, you do well in cross-validation and may do well on the public leaderboard, but your predictions do not generalize well to new data.
What this team did was to submit a lot of predictions and only take the predictions that improved public leaderboard score. Then they'd add slightly random noise and try to submit again. They repeated this until they ranked nr. 1 and left quite a few other competitors scratching their heads: How did they do this? Did they find a perfect ML algorithm? Is there data leakage? Are they cheating?
When the private leaderboard was revealed, this team dropped around 2000 spots. Their good performance on the public leaderboard was purely artificial. The contest was valid and well-organized (this could happen on any Kaggle competition with little data). They did not receive a prize.
There are other benefits to ranking well on the public leaderboard: It helps with teaming up with other high-ranking competitors and you can market yourself (recruiters are pretty interested in the top 10, eventhough the competition has not ended yet.)
Here is a hypothetical problem: Please write a calculator program, that takes input from stdin, containing one arithmetic expression per line (integers and + - etc), and simplifies the expression, returning one integer per line. Here's a sample input/output pair!
Input
2+2
3+1*2
Output
4
5
And you're all clever, and you write the following program: # python
print "4\n5"
Congratulations, you pass my sample test case, but you have overfit your model to the sample data, and it will not generalize to other datasets for the same task.If this were a Kaggle contest, and there were a leaderboard on how well people are doing on the sample data, you'd have a perfect score on the leaderboard. But then when it's time to see who wins the contest, and I try your program against some other inputs, you'd get 0%. Ha ha, funny joke you played!
The goal is usually to make interesting and insightful models based on a sample dataset. Overfitting occurs roughly when you tailor your model too specifically to the sample data. For instance, a regression with 10000 parameters might fit the data really well, but only because we included so many useless parameters.
A poorly designed contest on Kaggle had a submission that deliberately overfit.
If you're curious, there are many ways to address overfit models (look up model selection or AIC)