A Look into Machine Learning's First Cheating Scandal
dswalter.github.io
dswalter.github.io
- The LSVRC "visual recognition" competition has a rule in place, that limits to twice per week how often each contestant can run their entries to the contest against the ImageNet dataset.
- The Baidu team run their tests much more often, claiming that they understood the limit was placed on a person basis and not per team.
- This rule is in place because more frequent submissions somehow distort the quality of the test, by switching its emphasis to overlearning (adapting too much to the specific data included in the test dataset) instead of true algorithmic advances, which can influence the whole machine learning discipline given this contest's high profile. ( * )
- As a result, the whole Baidu company has been banned from the competition for a year.
( * ) ("The Baidu team’s oversubmissions tilted the balance of forward progress on the LSVRC from algorithmic advances to hyperparameter optimization.")
This is what the Baidu team were probing.
Some (all ? ) Kaggle competition also have a daily submission limit to avoid this kind of cheating.
People can, and do, overfit their model to the public test set, but doing so does not improve their score on the private test set, so cheating is prevented even without the submission limit.
The submission limit helps ensure the leaderboard generated from the public test set stays close to the leaderboard generated by the private test set while the competition is running, so that you can get an idea about your standing. But participants know better than taking the public leaderboard too seriously.
Alternatively, require the participants to provide an API and call them, instead of them submitting things.
But if you can repeatedly run different models against the test set and you do get to see a score, you could do something like random parameter searching or other optimization ideas to tune your algorithm to be highly overfitted to the test set.
Another way to avoid this would be to develop several test sets that are roughly "equivalent" in terms of the distributional properties, and then randomly change the test set periodically, or change the test set right after the final submission deadline, to discourage people from pursuing overfitting.
Now imagine a student can take the test repeatedly over the space of a few days, and can use the score to reverse engineer the answers to the questions. They can put in random answers and note which ones cause the score to go up. Of course the real life SATs don't allow this, and they change up the questions to prevent this sort of cheating. If this were possible, our enterprising/cheating student can derive the complete answer key over time, noting the changes in scores for each run. And once they have the key, they can ace the test. No longer is it a test of their aptitude, but rather of their knowledge of the answer key.
This scandal is analogous to this albeit contrived example. With an ML testset, it's not possible to change the data because you want it standardized so you can evaluate improvements that new approaches may bring. It's the only way to have a meaningful yardstick to measure against. Thus, the only way to prevent such gaming is restricting multiple submissions, so that you can't do 'hyperparameter optimization' - i.e. overlearn on the testset.
That's why it's cheating - it's not a measure of how well your algorithm did, but rather on how well you reverse-engineered the answer key. It's a huge disservice to the field and the people who did this should be ashamed of themselves.
> "The key sentence here is, 'Please note that you cannot make more than 2 submissions per week.' It is our understanding that this is to say that one individual can at most upload twice per week. Otherwise, if the limit was set for a team, the proper text should be 'your team' instead," Wu wrote.
I wonder whether, to a native Chinese speaker, this really does sound like it's talking about individual people, and saying "you" when one means "your team" seems really bizarre. Can any Chinese speakers weigh in?
(Even stipulating this, the affair still sounds more like malice than incompetence on the part of Wu.)
https://en.wikipedia.org/wiki/English_personal_pronouns#Arch...
Chinese, however, still uses both singular and plural second-person pronouns; 你 (nǐ, lit. you) and 你們 (nǐmen, lit. you all) respectively.
https://en.wikipedia.org/wiki/Chinese_pronouns#Personal_pron...
Note that the above is drastically simplified and I recommend reading the links for more information.
However, the matter of personal pronouns is such a fundamental aspect of grammar (at least in Indo-European languages) that it was literally the first piece of grammar introduced when we studied French. (It may not have been the first chapter, since the first chapter may have been limited to stock phrases and some pronunciation guides; it's been way too long). I find it highly implausible that any foreign speaker that has a working proficiency of English could somehow think that "you" is not used for both individuals and groups.
This sounds like they are trying to cover their malicious intent with an "it's better to ask for forgiveness than permission" kind of trick.
This explains the need to drill home the train-test idea from the last post. I hadn't thought about this before but multiple submissions do amount to multiple peeks at your held-out test set, which is a huge ML no-no.
I don't know much about LSVRC, but doesn't the way Kaggle work prevent this? AFAIR you get a "public" test-score which is used for the leaderboards, but once the deadline for submissions is up, each submission is evaluated on a held-out test set giving you a "private" score. Now that I think about it, I'm not sure how that works, I guess the accuracy they show you as your public score is only on part of the submitted rows? Regardless of how that's done, could the LSVRC organisers not do something similar?
One of the article's notes indicates that it most likely does, and that the Baidu team wilfully got around that limit:
> Members of the Baidu team had to create multiple logins in order to circumvent the “two submissions per week” rule
The absolute number of submissions is somewhat arbitrary, e.g. why are 40 okay, but 200 are not?
Without that, it becomes possible to pre-tune the algorithm to more closely match or better recognise the specific dataset ("hyperparameter optimisation") but not work any better in the general case, so the submitter does better on the specific challenge, but doesn't actually advance the field.
The limit is there to skew incentives towards algorithmic improvements, the specific rate doesn't really matter as long as it makes hyperparameter overfitting less convenient/efficient than algorithmic work.
What's the 'best practices' for the number of hidden layers in a CNN? 1 or 2 hidden layers?
And wouldn't video datasets be somewhat easier to analyse given the fact that you have multiple frames of the same object?
Think of it as using Hacker News posts to illustrate what posts are on topic. You could give someone a better impression if you showed them 10 posts instead of just 3, but if the 10 posts were all on the same subject then it wouldn't be of any additional utility. And if you accidentally include an off-topic post, then that user is going to have a mistaken impression of the post.
As for video... that's a long story.
While bigger usually means better, if for example there was a bias in how the data points were selected the data set can actually be worse than a smaller one where there was no such bias. Building a good data set really depends on the current understanding of the nature and difficulty of the task, both on the part of the scientific field and the researchers working on the data set.
The very same problem exists in the medical domain. How do you select the right patients to evaluate for a treatment? How do you know that there is not a specific genetic, racial, gender, etc. trait that leads to side effects? The answer is of course that you can not know with absolute certainty unless you include every human on the planet (and then there is the question as-of-yet unborn humans), but you can use your experience and medical knowledge to try to select patients (your data set) for the medical trial to minimise the possible impact of as-of-yet unknown side effects.
Didn't people claim "cheating" back when the first compilers started doing data flow analysis too?
Yes, who cares? The bigger picture is advancing the field, not scoring some bigger number in some artificial environment.
> This was about as cheating as it gets.
Like I said, so was data flow analysis originally.
--
The point is to not just dismiss this as "cheating" but take a closer look at how the current benchmark is flawed and how this sort of shortcut might be useful.
The point is to take a closer look at the current benchmarks and potentially come up with better ones.
https://en.wikipedia.org/wiki/Overfitting
Say the test data set, by chance, has 1% more dogs than the training set. By tweaking your algorithm to guess "dog" an extra 1% of the time you may be able to get .1% increase in your success rate. Also, that tweaking isn't a manual process it's part of the training process based on feedback from the results of your test set score. The improvement is enough to "win", but it's not because your algorithm is better.
That goes along with the argument that the data set is at its end of life. Maybe we're at the point that gaming the system is the only way to eek out the .1% needed to "win". In that case it's time to move on to a tougher test.
EDIT: I'd like to point out that the ImageNet competition is continually on top of the "time to make it more challenging" aspect. They introduced localization in 2011 (identifying not just what, but where items are). The 2015 competition includes, for the first time, recognition and localization tests in video clips.
function solve(letter) {
return {B: 11, H: 5, M: 11}[letter];
}
machine learning 4Head . Of course, the point is to solve for examples you haven't seen yet, so this "solution" isn't amusing anyone. That said, overfitting commonly happens even without trying to overfit, and people use techniques to minimize it.The parallel would be more like adding a detector for certain benchmarks in your compiler, and outputting hand tuned assembly for that case ... except even worse than that, because you'd have to implement the detector and assembly generation in such a way that it made your compiler behave worse on general input.
But what we're talking about here is much, much worse from a design point of view. It specializing your system for the benchmark, in such a way that you are actually making it worse in general.