InstaHide Disappointingly Wins Bell Labs Prize, 2nd Place
nicholas.carlini.com
nicholas.carlini.com
ML privacy isn't at all in my wheelhouse, but I did do computer security research in school. The author makes a good point that researchers need to be ready for their methods to be broken. More importantly though, researchers have an ethical responsibility to publicize when a vulnerability in their work is discovered, so it's not used improperly.
I think this is something that's very clear in the security community, but isn't the norm in the machine learning community. Hopefully privacy researchers in ML take a few notes from their colleagues that have been doing computer security sooner rather than later.
Finally, it's pretty disappointing to see how far Bell Labs has fallen. I won't spoil the article, but the quote from the Bell Labs judge is telling.
https://arxiv.org/pdf/2002.08347.pdf https://arxiv.org/pdf/1902.02322.pdf https://arxiv.org/pdf/1804.03286.pdf
He definitely doesn't mince words, but it's hard to be mad when he always makes very valid points. And I say that as a secondary author on one the criticized methods (not InstaHide).
There is a lot of machine learning snake oil out there, even in the most prestigious journals. The issue is that most people wouldn't know about it until they actually try to reproduce those results (if they can at all in the first place).
It is a huge issue now that machine learning has made its way into every field of research, while almost nobody has adequate training to even gauge what is and isn't questionable. Even the grad students who end up doing all the "hard work" usually don't have much of an idea of what exactly they hacked together and how to access its validity.
The crazy thing is that, despite all of this, you can still get ahead if you have know just slightly more buzzwords than others and have a paper or two (don't have to be first author) to show for. At the end of the day, I suppose all of this is nothing new (machine learning practitioners are the new SEO consultants?), and I'm just venting my frustration for not being able to accept this is how things work.
If we want to continue the snake oil salesman analogy, SEO consultants were soothsayers, attempting to divine meaning from The Algorithm of Great Google.
Now the apothecaries have begun to legitimize their potion making into tonics, but there is still room for a charlatan or two to push hodgepodge concoctions onto the unsuspecting Product Manager or CTO.
I'm also deep in that industry, and can confirm this is the case. Even for research coming out of the biggest names: FB, Google, etc.
The problem of "story over substance" has been observed by many; it even has a name in the context of ML, the "Mummy Effect" [0].
Like you, I'm segueing out of "AI" because of the unfavourable signal-to-noise ratio. Frustrating indeed, but life is too short to try to outcompete pocketless BS generators, or clean up their mess on my dime.
[0] https://rare-technologies.com/mummy-effect-bridging-gap-betw...
> In the award ceramony [sic], the Bell Labs researcher presenting the award explicitly said he doesn't understand how InstaHide is secure, but, and I quote, “it works nonetheless”. No! It does not.
Bell Labs used to have a wonderful reputation. Not so much anymore, I suppose.
That's not to say that M&A automatically (or significantly dilutes) the value of the research arm of a commercial organization.
Rather, every few years, you have to re-evaluate what they're actually committed to.
It's a small thing, but I'm really grateful that he very specifically aimed at the senior authors here. While I'm not sure that the first author can be completely absolved of responsibility (due to the simple fact that they are listed as an author), the power dynamics in academia are absurdly skewed and it's nice to see a tacit acknowledgement of that in an otherwise (justifiably) scathing critique.
Randomizing the sign is the same as eradicating 1 single bit out of however many were used to represent each sample. No matter how random it looks, there is no way removing 1/32 or 1/64 of the information in the image offers any actual security for something where samples are as highly correlated as in image data.
And using a non-cryptographic PRNG for security purposes... I'd really, really expect more from anybody working on privacy research.
I think this is the most abhorrent aspect: it's not in the spirit of science to arbitrarily delay findings
Anyway, why delaying results should be considered the worst transgression? I think dishonesty is the biggest problem here.
I'm not sure what it is about one-time-pads (and generalizations of OTP like shamir secret-sharing) that cause it to so reliably show up in unsound work. Something about these perfectly valid (in the right context) tools is very moth-to-flamey.
https://twitter.com/prfsanjeevarora/status/13374473410585681...
http://web.archive.org/web/20201211174253/https://hazelsuko0...
Is this unique to, or particularly common in, machine learning?
If so, that is scary - what on earth is going on?
Science was never meant to happen here, this is just plain exercise in dishonesty.
"Bad science" is a real thing which happens when somebody tries to "do science" but fails.
> Ben Goldacre’s wise and witty bestseller, shortlisted for the Samuel Johnson Prize, lifts the lid on quack doctors, flaky statistics, scaremongering journalists and evil pharmaceutical corporations.
> Full spleen and satire, Ben Goldacre takes us on a hilarious, invigorating and ultimately alarming journey through the bad science we are fed daily by hacks and quacks.
I find it useful to have separate terms for when somebody is incompetent (bad science) and when somebody intentionally misleads (quackery).
Edit: And I completely agree with author about Paper in the field need to do better.
Otherwise, it seems my assessment would hold. A ranking is relative, so doesn't tell you much of absolute progress towards the goal.