Maybe they need a "any score lower than X will be considered as a bast score", or score in other aspects of the solution (like complexity)
Or just have a lot of data for scoring purposes.
Maybe they need a "any score lower than X will be considered as a bast score", or score in other aspects of the solution (like complexity)
Or just have a lot of data for scoring purposes.
Here is the private leaderboard, which I believe represents the actual score of the contest, where they are in place 1,962: https://www.kaggle.com/c/restaurant-revenue-prediction/leade...
That's, ah, not in the prize range, to put it lightly.
There's no problem with the contests here. It's just a particularly vivid, and entertaining to humans, demonstration of the well-known problem with overfitting.
If I am reading this correctly, the green numbers to the left of the teams in the private leaderboard represent that team's position relative to the public leaderboard. Note they're quite large... the 9th place team in the final scores jumped 1,058 positions! That's nearly half the field. If this is typical, and I have no idea (reply & let me know!), the public leaderboards are basically a joke.
(there seems to be simple ways for kaggle to solve this: put a delay on the score generation, limit number of submissions, limit precision of scores on the public leaderboard)
The published dataset gives contestants an idea of how good they are, but prevents gaming the system by overfitting, because the contest is actually scored with the second unknown dataset.
A more in depth explanation on the use of three datasets for model building can be found here: http://stats.stackexchange.com/questions/19048/what-is-the-d...
There is a limit on the number of submissions. Usually 5 or less a day. Also precision of scores is visibly limited, but for final ranking all decimals count.
This forum post discusses competition variance and poses a metric to quantify "leaderboard shake-up": https://www.kaggle.com/c/liberty-mutual-fire-peril/forums/t/...
The Public Leaderboards are very helpful though! When you have setup a solid local cross-validation pipeline, and the public leaderboard agrees with your local evaluation, then you can try a lot more algorithms and parameters, without using any submission. Especially when working in teams this is important as you may have only 1 submission every 2 days.
Also, the more advanced Kagglers can use leaderboard feedback to increase model accuracy: Cluster the data sets with objective measures. Apply a modifier (restaurants from this region get 0.95 x previous prediction) and look at the result. If the split between public and private is random, and your clustering is objective, then an improvement on public leaderboard should reflect in private leaderboard.
That said it'd obviously be nice to build leaderboards that are robust to this type of gaming. There's been some recent work, mostly driven by Moritz Hardt AFAIK (briefly referenced in the submission), using ideas from differential privacy to create leaderboards that leak only a bounded amount of information about the validation set: http://arxiv.org/abs/1502.04585.