A Kaggle Grandmaster cheated in $25k AI contest with hidden code
theregister.co.uk
theregister.co.uk
I feel as though comments are less informative, more judgemental, and chalk full of emotions.
...as if a younger crowd of users has started using it.
Have the older-wiser HNers started using another site?
In my family if someone cheated they got called a cheater and suffered consequences. At least, they would have, if someone did something like that. But my parents didn’t raise mendacious villains.
Look at this crap on Twitter:
“Everyone makes mistakes. Thank you for the apology”.
“Kagglers will still love to have you back”
“It's great that you realize your mistakes. Looking forward to see your comeback with more cool DS solutions and ethics than before.”
“Thanks for doing this. It's okay to make errors in judgement, we've all been there to varying degrees. Y'all be gonna be fine.1!”
Those are the worst. I’m not so crazy about these below, either, although there’s just a hint of steel in them, at least:
“I’m glad to see that you had a change of heart after sleeping on it and that you will be returning the prize money. I hope you will consider donating to or volunteering at a local animal shelter as well. Atonement here is more than returning the money and apologizing.”
“I hope this can be used as a teaching moment as well. Many people clearly look up to you because of your work. What can we learn from this? Something to ponder in the days to come.”
He was publicly shamed, banned from Kaggle and he lost his job (and I guess he's basically unhireable right now).
I wonder what would be your family reaction, if these things don't sound like consequences to you.
A lifetime ban would not be out of the place here, considering it went on to win the competition with no admission until caught.
Some headhunters specifically try to identify high-performing sociopaths for top management positions.
[Citation needed]
You picked a random sample of ratio'd comments from a platform which is big on performative wokeness to strengthen your opinion, which still fails to explain how Pleskov managed to zugzwang his way into pulling off a Kobayashi Maru style move and the failure of the platform to monitor such abuses. Only time will tell, whether or not he has avoided being a part of the Dark Triad; without condoning his behaviour ─ what more can he do to atone for his sins?
So what's the problem?
The prize was $10,000, not "fun".
More importantly, Kagggle does competitions across dozens of industries. If a culture of getting better at hiding your cheat pervades the platform, that could impact finance, transportation, and medical research. In those scenarios, lives would either be threatened or at least subject to sub-optimal systems.
I compare this to my favorite sport, Formula 1 racing. In F1, teams of engineers with nearly unlimited budgets spend an absurd amount of effort doing everything in their power to bend the regulations (the "formula") to squeeze out some extra advantage.
For example, in this past season Ferrari was suddenly outperforming the pack (and their own recent performance) and it was clear something had changed on the car, they had power in places they didn't before. What finally came down is a clarification of the rules around fuel-rate metering, without directly calling out Ferrari. After the clarification, Ferrari power was back where it used to be. Nothing more was said of the matter by the FIA.
What we all _think_ happened is that Ferrari, knowing the fuel rate meters ran at 10kHz, discovered they could pulse their fuel pump so that the low-end of the flow rate cycle happened during that sampling interval. This means they could increase their overall fuel rate beyond what was technically allowed, due to how that technical requirement was being measured on the car (and reported back to the FIA).
Is it in keeping with the spirit of the rules? Of course not! Does it make for an interesting engineering puzzle on top of an already-exciting sport? Sure does!
Clearly I'm in the minority here, but I think this sort of problem-solving approach can be useful. If you're looking to compete against a field of entrants who are all looking for obvious and well-understood approaches to solving the problem at hand, I think sometimes the best solution to stand out is to look where the other teams aren't looking.
To use your F1 analogy, this isn't the equivalent of tweaking the cars in whatever way possible is within the rules. This is the equivalent of completely cutting across the grass and bypassing 90% of the track, which is indeed illegal and would get you penalized.
Jesus fucking christ. He fucking ran MD5 on some shit he pulled down from a web crawler.
Over-training a model on the validation set would be a lot more "brilliant", and even that is a dumb script kiddie level hack. Maybe finding an algorithm that computes weights s.t. the preimage of the training algorithm on the training set matches the result of training with truly random weights using the validation set. That could be a "clever hack". And even then _brilliant_ would be.... a real fucking stretch.
Of course, we're all human, and he's come clean, but his actions potentially had a negative effect on the non-profit and the animals it places; and competing talents were denied their rightful places.
This comment isn't about condemning him or anything, just let's be honest about what happened here; it wasn't ok, or just system-gaming caught out.
Absolutely. What are the odds he (and/or others in his team?) did this just for this one competition? It's possible, but unlikely he/they invented this (or other) code hiding technique(s) just for this single occasion.
Also, the employer did the right thing and kick him out. Now they only have to scrutinize the last few months of his work instead of looking over his shoulders the next few years.
https://news.ycombinator.com/item?id=22045696
I posted there, same self-addressed question, that I cannot figure out the answer to...
It seems that intensives to cheat, and environment where 'means justify the ways' -- are overpowering.
For people who are naturally gifted, successful at young age -- why cheat?
Was this historically, always like this?
These insensitive to cheat, to gain unfair advantage, to treat life opportunities without any 'honor code' just seem to be so pervasive now, it seems.
There is a cheating scandal every other week involving most prestigious institutions, competitions, and so on.
These incentives to cheat, basically destroy from inside our commercial model, academics, judicial system, political system and probably military too.
This also creates a new type of powerful currency, and therefore the 'billionaires' in that currency have infinite power -- and that currency is 'dirt on somebody'.
Dirt on somebody who cheated before -- forever makes the cheaters into tools of injustice.
---
Public shaming is reactive, there we need something more proactive at various points. There are needs to be incentives for work verification, as an example.
I also think it is unfortunate but at least civil/commercial law in many countries is pretty much riddled with 'more expensive lawyers produce better results'. And it skews society into basically thinking 'anything goes, really. means justify the ways, and cheating something one can get away with'
On topic: My university offers "Ethics for Nerds [=CompSci]" lecture. The lecturers have degrees in both CS and Philosophy, and the stated goal is to make compsci students more aware of ethical implications - plus giving them some tools/thinking to assess these implications.
I also took an ethics course, but in there it was mostly about AI impacts on society, and what will happen when people loose jobs that will be automated away...
I should keep up to date on it, as mine was many years ago.
After all, being a computer programmer has to be more than about VC funding, mobile apps, functional programming, AI and kubernetes. :-)
To me, the seeming prevalence of cheating through out the society, and its tacit encouragement, by lack of effective proactive and reactive deterrence - is a cultural, as we well legislative problem.
Many are used to being the best. There’s an old saying that 90% of students aren’t in the top 10% at Harvard. When your entire life, you’ve been the smartest kid in your class, it is a tough adjustment when that is no longer true. So they try to find ways to get that top rank again.
How can you ensure your kid can handle this? Make sure they are exposed to situations where they aren’t the best and reward them for giving their honest best.
In this case, the cheat was obvious, but how do you detection enrollment cheats or totally bought research projects? It's not even funny how little you can do here.
Even plain and obvious plagiarism is being missed.
Of course. 20$ bills left on the floor are bound to be picked by someone even if most won't.
..At least as long as the expexted punishment comes down to less than the value of said bills.
And honestly, there should be no guilt in that. Now, had there been someone going around saying "hey, I had this money here and the wind blew it out of my hand, have you seen it?" - and they named the denomination or something that made you know that the money you picked up belonged to them - and you didn't say "why yes, I found this over there; here you go!" and handed it back...well, you should feel guilty knowing you had taken their money - even if they didn't see you do it.
But if nobody is actively looking for it, then heck - what can you do, and why feel guilt over it? Sure - you could have left it lying there, and if everyone did that, maybe somebody could retrace their steps to find where they dropped it?
Or for all you know, it was dropped, not noticed, and the wind or a passing vehicle picked it up and flung it hither and thither and even if the person knew they had dropped it, they had no way of finding it. You could feel guilt - but now you are feeling guilty over something that is virtually unknown to you; you don't even know if the person who lost it even knows they did, or if they even care...
But again - it's a different thing if you see it happening. As a kid (I was probably 12 years old), I was once in a bowling alley when I was walking behind a man, and I noticed a large amount of money (bills) fall out of his back pocket, and he didn't even notice one bit. We're talking several 50 dollar bills. I saw it. I saw him walk on. I went over to pick it up. It was a lot of money...
I grabbed it all and ran to him, "Mr! You dropped all of this!" and handed it back to him; he was super thankful. I was glad I did it, because I know most other kids wouldn't have done that - heck, there are adults who wouldn't.
But I had seen the money fall out of his pocket - I knew it was his. Had I kept it, I know I would have been stealing from him, and would have felt guilty. It was the right thing to return it without any expectation. To this day (decades later) I am glad I did what I did - that very well could have been his paycheck for all I know.
Now - had (somehow) that pile of money been there and I saw it, and nobody else had claimed it? Well - even today I'd probably take it to the "lost and found" - somebody loses that kind of money (or item - like a wallet, phone, etc) - that's the right thing to do. Only if nobody ever claims it, and it's returned to you as "unclaimed" - should you take it. Even then, maybe it would be better to donate it to a good cause than to keep it.
Definitely an ethical and moral dilemma - and here I don't know if I've just refuted myself, or whether I have helped anything for you, or if I've just made myself question my own failings or faults. Probably a bit of everything - and maybe that's for the best. So thank you for your post - I think.
In this competition, the training code was run on Kaggle's system, so you'd still need to smuggle in the extra data.
You've got the testing set. Create random HPs and tune them to fit. The way they cheated is stupid.
And the way the testing set can be obtained is silly.
Considering the guy was smart (he is kaggle grandmaster), I would really like to know what prevented him from training on the scraped data, and what motivated him to obfuscate the known sample lookup.
Maybe there's some technicality they made it impossible to tune the model on the additional scraped training data.
" Thus concluded one of the most unusual of my adventures and voyages. Notwithstanding all the hardship and pain it had occasioned me, I was glad of the outcome, since it restored my faith, shaken by corrupt cosmic officeholders, in the natural decency of electronic brains. Yes, it’s comforting to know, when you think about it, that only man can be a bastard. "
(source: The Star Diaries, Stanisław Lem)
The goal of this competition is to build a system (using ML or not) which is useful for predicting how quickly pets will be adopted. Any information used during the competition should be realistically available at inference time for future predictions... clearly, the expected answer cannot be available at the time you're trying to predict when a pet will be adopted.
If one can excuse scraping the data to build a better model, I can't see how one can excuse this.
He wasn't really making predictions at all, just looking up the answer.
And the goal really is a prediction system, not something that looks up previously known answers.
One test a good of ML competition: Can it be solved by simply hiring lots of humans to make predictions without incurring significantly more costs than the prize money?
What value to the organizers, to society or to whatever are you imagining coming out of a free-for-all style competition?
I think the organizers now imagine that the result would identifying good, generic prediction algorithms along with identifying good AI programmers capable of producing general prediction algorithms.
It seems like the contest framework already has become a bit problematic through context winners just being good at contests and not otherwise achieving anything.
But what are you thinking of? There are already hacking competitions btw.
How? According to the article his model without the cheat rated at ~100th place, and the article mentions him cheating the same way before (by scraping Quora for some Quora related competition).
> These predictions would be used to optimize and tweak future critters' profiles so that they are adopted as soon as possible.
Sorry but, how is this useful? You can't just change the age of an animal to make it more likely to be adopted. The profile is meant to be an accurate representation of the animal so people know what they're getting. What exactly was the algorithm meant to achieve aside from being a predictor?
I'm not sure how this website manages "inventory", but they might have similar problems.
This contest sounds ridiculous. It sounds like an attempt to get in on that AI gravy but do so with some sort of feel-good element. Only there is no feel good to it, and the basic premise seems outlandish.
They do, though. That's how they are able to limit the number of animals they have at any given time.
And just to provide the full picture, most no-kill shelters of course have scenarios where they euthanize -- violent animals, sick animals, etc -- but they don't need a neural network to accomplish this.
This is all neither here nor there, as the contest had positively nothing to do with any of this. Instead they wanted to determine the most adoptable traits so they could adjust the less adoptable traits with the more adoptable traits: The poodle goes through the hair straightener and gets a blonde hair color treatment (clearly I am being satirical) to make it more like a lab, for instance.
A limited number of parameters can be genuinely altered - a better photo can be taken, and vaccinations can be administered for example.
Animal rehoming centres have to balance throughput with cost; reduced per pet costs mean that they are able to expand or support more complex cases.
Whilst keeping a pet in a "space" and feeding it does cost, this cost can easily be significantly less than vaccinating, particularly as some vaccines require a few days hold post vaccine. Similarly if the vaccine will not alter a pets rehoming chance, then it is an unnecessary cost.
Pictures may be more easily applied as a tighter feedback loop (of the 5, use the 3rd) however they may also indicate other issues that could be addressed (over / underweight, coat damage, etc.) and addressing those issues have costs to balance and predict.
I think one of the biggest skills to have in the ML space, is knowing what is worth training a model to know, and what isn't. Just like in engineering, the most successful products are those that solve a real world problem, no matter how elegantly the others might have been made.
But I would be really interested to see if it really has an effect, and if that effect can be sustained.
I think this is the closest to the truth, but that Google spent the money.
Google wants to show usefulness of ML. Marketing person comes up with competition and backs it. Pet adoption org just has to provide some data and says why not? At the very least, it is free publicity for minimal effort.
> In this competition, the training code was run on Kaggle's system, so you'd still need to smuggle in the extra data.
The question then becomes, how do you smuggle in the data? This is a much more interesting discussion than pontificating about the ethics of Pleskov's actions. In particular, a better understanding of this problem could have ramifications for how Kaggle could combat hacks of this variety. (By contrast, "shame on him" and "aww but he's a nice guy" are both useless, except perhaps as a form of virtue signalling).
It's essentially a cryptography problem. Does anyone know if this has been widely studied?
Solving this would require changing the contest so that it comes with algorithm and instructions only (no data files, entropy checks) and is trained by contest operators.
One approach would be to write some handcrafted rules/features that look likecthey were plausibly a priori, but have the effect of memorizing the scraped data. (I don't know if this is actually possible.)
User KaoruAoiShiho is probably right that any approach along would look out of place. Coupled with the fact that its removal would massively boost accuracy, it's hard to imagine how this would get past a curious reviewer.
Perhaps peer review should be a component of the Kaggle process.
Your point of view is outdated ;-)
In recent ML competitions, participants do submit code that is run on a held-out dataset - as was the case in the PetFinder.my challenge in question here.
Most competition platforms are migrating to this format, as otherwise you can just label by hand as you said.
Note that this competition went even further: not only was the evaluation code run on Kaggle, the training code was also run there. This means that you couldn't even train a gigantic model then submit it: your model had to be trainable within well defined time and resource constraints, which is a great way to level the playing field.
Of course, there's still some unfairness as people with more resources can try out more solutions before submitting a model to be trained on the platform. No platform has a solution for this yet!
The "cheater" in this competition is both a world-class data scientist and a reverse engineering hacker. Heck, people used to write papers about how they crawled the ground truth. Now they see their name in the newspaper.
This is true of academic contests in general, btw, even without cheating. They stop being interesting/fun/good signals as soon as people start treating them as an independent skill set. Comparing performance on chess games for the first N games between two new players might be a good signal for some general intellectual capabilities. Comparing experienced players against one another is mostly just testing who's spent more time learning about chess.
Has anyone cheated at Kaggle/similar using that approach?
That said, the site is a fantastic resource for datasets. Lots of fantastic data uploaded both from old competitions and by the community.
The whole point is to match abandoned animals with suitable homes. Adoption speed seems secondary to post adoption measures of adopter satisfaction and adoptee welfare.
Also what you are writing are extremely sparse and weak signals compared to adoption speed.
Like in business and tech, you want a high as posible throughput / inventory turnover rate - but obviously with constraints.
I'm not even certain of how that could be objectively measured for children, let alone pets.
They did not implement the winning solution as is, but a lot of good stuff came from the winning teams, including techniques (such as SVD) that are in use as of today (or maybe a few years back).
> just a nonsense engine
No, it is safe to assume the recommendation engine of Netflix is close to the state-of-the-art. A lot of money and talent went into it.
Shortly after that contest finished Netflix removed the five star rating system, and dramatically subdued the recommendation engine. Now the vast majority of the content surfacing is universal beyond some small category filtering (e.g. You like crime dramas and horrors so here's a bunch of the most popular stuff from those categories).
Now they have a "match" rating that I have not met a single person who finds useful (it is almost at the point of farce and seems more like a randomization engine).
A lot of money and talent went into something that Netflix clearly decided just wasn't important or useful enough for them. Now they just push Don't F*ck With Cats on everyone.
How is that different from "LeetCode PhD" doing "DP Hard" for 6 months to get into FAANG?
Now it really illustrated to me how data science is still an incipient field. Even the basic Titanic example has a lot of fake solution and a bit of "cheating" (overfitting).
And don't get me started on the multitude of tutorials and examples that only show the loss decreasing graph but the actual results are not evaluated (and thus might have some glaring defects).
They're essentially letting an algorithm decide which dogs have the best chance of adoption and which to euthanize, aren't they?
I mean, you're right, you could just run the tool, find the ones that are unlikely to get adopted and euthanize them, but I don't think there's any reason to believe that's actually their intention.
[1]:https://www.kaggle.com/c/petfinder-adoption-prediction/overv...
It's not hard to jump from this a system that would use your facebook file to grant or deny medical health coverage or a similar life-and-death thing. Shift the Overton-window a little, push is a few notches further, and you're re-inventing phrenology with deep learning...
This is bone-chilling! I mean the fact that so many people overlook this...
Of course this is highly problematic when applied to humans. But to be clear, some humans are already judged by models (transparent ones so far) in live and death decisions. That's how organ transplant decisions are taken. The candidates with the best prospects get the scarce organs. Similarily, when looking at populations, decisions that protect many are often taken in full knowledge of the danger to a few. Vaccination for example.
Resource optimization problems mix badly with absolute morals.
https://twitter.com/ppleskov/status/1215983188876709888?s=19
The person was originally employed at h2o.ai and as a consequence of this was fired. Not sure if that was completely appropriate.
Wasn't this a personal participation? Or are there "company teams" on Kaggle ?
In my opinion, keeping him onboard compromises h2o.ai's image as a trustworthy AI company. To many people, AI is magic, so the last thing you want is the impression that the magic is really a con.
He has only been at h20 for a few months.
He was almost certainly hired at least in part because of his participation in the contest.
If you admitted to cheating on your degree, and your university found out and revoked your degree, then you'd probably expect to get fired from your job if you no longer have the qualification you said you did. Being a top 10 kagglester is a qualification (and a highly prestigious one at that).
Put down the pitchfork and calm down. The guy is a raging a* and a cheater, but a criminal he is not. He lost his job, was publicly shamed and his reputation is tarnished basically forever. I think he got enough coming his way.
"Sorry, you are likely a cheater, we logged you out and banned forever."
Competive ML model grading with a common training set using unseen data.
Cheat was to scrape the data that would then be used as unseen by the organisers. Unseen is now seen for this model. Then, instead of training the model with the "unseen" data, which would have been cheating and an advantage it apparently wasn't enough of an advantage so they hard code 10% of the cases to boost metrics and win.
Having more data to train your model is google & facebricks competitive advantage. Their attempts to use that advantage for something actually useful to society rather than just as a method of selling ads seems to have been a complete bust so far. If that is wrong you know better, please link us up.
I'm suspicious that their predictive power to sell ads actually works for the people who buy those ads but I guess we aren't likely to know for sure. I do wonder "who dominated their industry segment in sales by being an early adopter of google ads" I don't know anyone. It's not a great metric but what else do we have?
It's looking up the answers in the grading sheet while taking an exam.
You have a model that you have trained on some provided data - the training set. You give kaggle this model. Kaggle grades your model on some different data your model has never seen. The better your model classifies this data it hasn't seen the higher it scores and the more money you win.
So again if you trained your model (code) on a training set that you have illegally obtained. That is broken the rules of the competition to get the additional data. Data that is also going to be used for competition verification and grading, then you have cheated. Doubly so if you just hardcode outputs to boost your score, which they did here.
They took the official training set. Said, we need more. And scraped websites to get a bigger, illegal training set. This is against the rules and is cheating. They got caught.
It really is equivalent to looking up the answer sheet while taking an exam.
So yes you absolutely /can/ submit a model trained on your own private data set even if what you submit is a model code that will be re-trained. Even if "the training set" is different to the provided - you still have that scraped data so you can slice it up with the provided training set so that any selected training set does well against the rest. Now the overfit you've just carefully engineered should win against the honest models unless you suck, right? It's kind of risible that they had to go further and hard code certain results, don't you think? Perhaps if they still couldn't win with scraped, illegal additional data then everyone else had illegal data too? Perhaps Kaggle is not a good indicator of how good ML techniques are in practise? Perhaps Kaggle systematically overstates ML effectiveness due to this kind of uncaught cheating in many of their competitions? I bet kaggle won't look too hard at that.
I'm by no means an expert in ML, but my understanding is there's some code that is run to train the modal. I meant that by "training code". My regrets if my terminology was unclear.
> They took the official training set. Said, we need more. And scraped websites to get a bigger, illegal training set. This is against the rules and is cheating. They got caught.
No, this is wrong
I suggest you look at one of the other comment where people have explained why this is wrong. They did a better job than me.
https://news.ycombinator.com/item?id=22124760
Similar with school: how many times you, excellent grades holder, was cheating because no one scrutinizes high grade holders too much?
Accumulated integer counts that represent social value and/or standing encourage destructive behavior.
Perhaps someday we'll get a week without them.
(I’d love to be surprised, though.)
You can get the same effect with a Tampermonkey script I wrote: https://news.ycombinator.com/item?id=14456200
It even scrambles the karma count in your profile, so there’s no opportunity for karma to affect you.
If "improving individual performance" means to you that you perform better at your job, you are already lost.
The effect on parents seems bigger, which is a good thing, I suppose, since parents then give their kids to some positive feedback.
I participated in some of them before due to the visibility. People from your class may know you but others, not so much. In the absence of time and other signals, management would often pick someone who is highlighted to represent something.
A lot of big people at those events generally don't mind giving contact information to young kids asking for help as much.
From there, you can effectively build a nice list of people that you can leverage in future for your resume or career.
It's like free consultation you otherwise would need to buy later.
Jeff Atwood on Coding Horror. (He can't be the first person to have said this, though.)