Grand Prize Awarded To BellKor Pragmatic Chaos
netflixprize.com
netflixprize.com
"The data set of more than 100 million entries will include information about renters’ ages, gender, ZIP codes..." It's enough to identify 87% of the people, apparently: http://www.freedom-to-tinker.com/blog/paul/netflixs-impendin...
I hadn't realized that someone identified some of the raters in the previous contest with their imdb ratings: http://www.cs.utexas.edu/~shmat/netflix-faq.html
Also, why this matters, page 44 of a research paper (PDF): http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006
the difference between the 'test' and 'quiz' sounds b.s. to me. at this rate, i know i won't even be contemplating netflix prize 2. at the very least, netflix owes it to the community to explain why they (at this point, seemingly corruptly) made the decision they did.
i suppose they say they'll post the final 'test subset' scores on the leaderboard. it doesn't appear to be a very 'open' contest if they don't actually publish exactly what the test subset is and how the rankings are determined. when it's that close, everything should be shown to be exactly done within the rules. otherwise, they really risk delegitimizing the whole contest, imho.
As for why the intermittent leaderboard shows results based on the public dataset: using a specially crafted algorithm, you could leak information about hidden data through any feedback you get from the test environment. Such information could be used to create an aforementioned algorithm that does well in this specific case.
see also: http://en.wikipedia.org/wiki/Overfitting
i'm not saying netflix is unfair or corrupt. i hope not. but i think they owe it to the contestants to explain exactly why they chose the #2 entry on the leaderboard over the #1 entry. and not just (though i appreciate people's comments) because of overfitting. how is the private 'test' subset chosen? i have no affiliation with this contest or any of the teams in any way, but how would you like to work on this for years only to be told that even though you beat everyone in the quiz, in the private 'test', you lost.
Then the ones over 10% were run against the final test set and the best winner determined from that.
Obviously releasing specific details about the final test set doesn't work because teams can fit to that :)
Now, on the front of selecting the winners before the contest has even started, and telling them about how to bias their algorithm so it will do better on the final--I can't really help you. I'd just hope that the programming teams would be honorable enough to mention if one of them were offered such a thing.
Also, don't brush overfitting aside. In the (paraphrased) words of one machine learning researcher, "life is a battle against entropy. In the same way, machine learning research is a battle against overfitting." Any data whose test results you use to select your algorithm or to adjust your algorithm's parameters is not properly considered part of the test set; after your optimization that data will provide an optimistically-biased estimate of your true error. Since competitors could get regular updates on their performance on the quiz dataset, one must assume that they were attempting to optimize this performance, and so quiz set performance was not be a good estimate of their true error. You can only get a good estimate of the true error of a method by testing against data that has played no part in its development.
Yeah....what happens when the checksums don't check out?
The "secret" data set is in the file "judging.txt" in the grand_prize download.