The auction that set off the race for AI supremacy
wired.com
wired.com
Woah, this came out of nowhere and it’s completely wrong. The problem isn’t that deep learning is picking up biases of the researchers, it’s that it picks up biases from the training data.
audible sigh
Come to think of it, isn’t that an interesting venue for GAN-esque methods to detect the relation of patterns falling into these categories of biases? Or is that recursive problem? If not, put me in the paper :-)
That's not strictly true. In a lot of cases you start out oblivious to biases in the data, and then when you evaluate the model you notice problems.
But your point about obliviousness to bias is exactly what I'm speaking towards. One might be oblivious to bias that aligns with your own biases, but notice bias that conflicts with it.
That's not a systematic way of tackling bias. I would rather have invested more in creating good benchmarks and norms.
The idea (as quoted) that models are routinely picking up biases directly from researchers is complete nonsense.
[0] https://twitter.com/sarahookr/status/1361373527861915648?s=2...
Between all these the degree of L1 regularization or the class weights are minor things. Most models will perform similarly given the same data. It's mostly the data that makes the difference.
But if you take "model" to mean the pure mathematical description without parameters or hyperparameters that need to be determined by experimentation, then I agree that optimizing the model on a dataset will not lead to bias against specific groups of humans unless the data used contains such a bias.
But this ties back to the original data problem, right? If you don't have enough training samples for (known or unknown) unknowns, your model is likely to be biased against them.
It is a known challenge to align the designed purpose of an algorithm with actual optimization metrics. For instance, recommendation systems may have the purpose of improving user experience, but if time-on-site metrics are used as the optimization function, there can be unexpected results.
Given that reducing bias while not giving up other desirable properties is a young and open research direction, researchers in general should not be faulted for using the current (imperfect) state of the art or for working on something that is not (yet) focused on bias.
Even humans need to know about swear words in order to consciously avoid using them, or need to learn about reproduction in order to avoid teenage pregnancies. Not knowing does not make us or the AI better.
For example, what GPT-3 needs is a "conscience", a separate model monitoring and rejecting harmful outputs. If I am not mistaken the demo is already displaying warnings when it goes off into weird places.
This is not only false, but in the context actually intentional misrepresentation. Most of the issues with the model was solved by introduction of hidden layers and backpropagation learning, which is at least in my opinion required knowledge in CS since at least early 90's and probably earlier (it is not clear when the idea was formulated in usable form, put the most cited publications are from late 80's, eg. Rumelhart, D., Hinton, G. & Williams, R. Learning representations by back-propagating errors. Nature 323, 533–536 (1986). https://doi.org/10.1038/323533a0).
On the other hand obviously more complex modern approaches to the "throw bunch of poorly understood linear algebra at he problem" problem have value and there is definitive generational shift in the current "AI-anti-winter" (for lack of better word), but still...
The best techniques we learned were MCMC and random forests, our computer vision was OpenCV and didn't work so well, and there wasn't any suggestion that buying a lot of GPUs and not bothering to understand the problem space would produce better results than our laptops.
Key Quote:
"Inevitably, the next bid wouldn’t arrive until a minute or two before the top of the hour, extending the auction just as it was on the verge of ending. The price climbed so high, Hinton shortened the bidding window from an hour to 30 minutes. The bids quickly climbed to $40 million, $41 million, $42 million, $43 million. “It feels like we’re in a movie,” he said. One evening, close to midnight, as the price hit $44 million, he suspended the bidding again. He needed some sleep." (So Google paid north of $40m to hire Hinton and his lab).
That all makes $2m a year for Ilya Sutskever, or $600k a year for some run-of-the-mill dude with a degree in AI seems like a bargain. The founders of Nuro got $40m each to leave Google and start their own company. Hopefully we get some more transparency about pay to bring up the whole field rather than just "AI." After all, at the time of writing software tends to have a higher accuracy rate than machine learning ...
“In the days before the auction, [Microsoft] complained that Google, its biggest rival and likeliest competitor in the auction, could eavesdrop on private messages and somehow game the bids. Hinton had raised the same possibility with his students, though he was less expressing a serious concern than making an arch comment on the vast and growing power of Google.”
NYTimes is positively in love with him, are a few other articles - https://www.nytimes.com/2017/11/28/technology/artificial-int... ... https://www.nytimes.com/2016/12/14/magazine/the-great-ai-awa...
He was on one of the first backprop papers, worked pretty consistently on neural nets though the AI winter, and was pretty instrumental in the renaissance.
Was he the only person working on this stuff? No, but I think we as a society like to tell ourselves a story about science and scientists. The narrative of the “lone scientist” wiling away their days in solitude until reaching a breakthrough probably hasn’t actually existed since the 19th century. Pretty much all significant discoveries nowadays are the product of intense collaboration and iterative progress. Science is a team sport, just one where 99% of the team happens to be invisible.
There are countless articles that romanticize entrepreneurs who were child chess champions, or rubik's cube solving geniuses, who dropped out of college, or were academics who were under-acknowledged, only to build it all by themselves and hit the jackpot. At least that's how the story often goes. It almost seems like a sort of modern mythology, one that taps into the American dream that so many yearn for.
Were the students or Hinton under any obligation to stay at DNN-research?
for example:
3 people x 5 years guaranteed loyalty = $45 mio
=> $3 mio per person per year, paid upfront
That's still high, but not that much higher than what Google already pays to top performers. Plus they get all the PR benefits, as we have seen.
Neural nets are incredibly good.
This stuff is already making massive impacts, and yet it’s still in the early stages. The meeting in the article was less than 10 years ago. Just imagine in 20 or 30 years from now.
With that said, I generally feel like my life has not been impacted greatly. Siri can barely tell me anything other than the first paragraph of a wikipedia page or today's weather. I don't search my photos often. I don't translate languages often, and I'm not in a back office in charge of moderation.
- Siri, Alexa, Google assistant all use deep learning for both understanding your speech and creating artificial voices
- Automatic translation
- Anything where you search for images or with images
- Virtual lenses in Snapchat/Messenger/etc, portrait mode and artificial depth of field in cell phone cameras
- Any recommendation system like YouTube, Spotify, or Tik Tok
- Tesla autopilot
- Image upscaling in video games
- Facebook/Google/anyone targeting you in ads
e.g. look at translation. It was not usable a decade ago
https://en.wikipedia.org/wiki/Vickrey_auction
This ignores phycological factors, and some minor quirks relating to minimum bid increases.
You could drive up the price by continuing to bid even when past your threshold, but only after an opponent also bids. But first, that does not get you any benefit, it just inconveniences a rival. And second, this will lead to a tie, which presumably has 50/50 odds of you ending up the winner.
I had a similar question about the backpack: they thought they could open it and find out what Baidu’s bidding strategy was. But that knowledge would have had no value to them. The outcome would not change.
If Baidu is currently winning at say $21M, Google might bid $22M if they know Microsoft will not bid in that round, but might bid $22.1M if they know Microsoft will bid $22M in that round. (Or $23.6M if they know Microsoft will bid $23.5M.)
Or suppose Microsoft is willing to go to exactly $22.5M, Baidu to exactly $23.0M, and Google to $40M. If Google can get that information, Google can bid $22.6M and beat both Baidu and Microsoft (if Baidu doesn’t bid this round, because they’re anticipating a $22M bid this round that they’ll cover at $23M next round).
Furthermore, execs are creatures of emotion and whimsy, rather than strictly rational operators. As the price soars higher, they see that everyone at the table values the thing highly, and are likely willing to increase their max bid as their sense of FOMO increases... Subincrements aren't going to matter much, except to ratchet up the feelings at the table even higher, faster.
Another surprising thing is that there were no good will clause, ie that they could terminate auction and refuse to work for the winner (which they effectively did).
Under independent private valuation (where the value of the object is completely idiosyncratic to you), your thought is (AFAIK) correct. Here think of art: ignoring resale value, how much someone else would pay for an object depends on how much he likes it, but doesn't have anything to do with how you should bid - your bid is a function of how much you like it. (To a first approximation - again the details of the auction format really matter.)
In a common value situation (where the value of the underlying object to any player is the same, but players _signals_ (an input to their beliefs) about that value contain an idiosyncratic component), you might learn something about the quality of your own signal by knowing something about the bids of other players in the auction. For example: are you about to overpay b/c you believe Hinton's time is worth a lot more than Google and Baidu believe it to be worth??
The phrase "the winner's curse" was invented for this problem in the context of bidding for offshore oil leases. The value to all players (the number of barrels of oil in the field) is the same, but they have different beliefs about what that number is. The winner's curse is that the highest bidder was generally the most optimistic about the value of the field. It's a "curse" b/c the optimism was frequently not justified.
What was actually being auctioned here is a little unclear - something like "the time and expertise of Hinton + two grad students", but it's plausible that Google, Microsoft and Baidu could have made equally good use of it. It reads more like a common value setting, that is. So knowing the bids can matter.
If you want to know more about auction theory (this comment may already be more than you want to know!), I can recommend Krishna's "Auction Theory" or Milgrom's "Putting Auction Theory to Work."
Of course in the end it didn't matter! Hinton subconsciously knew he wanted to join Goog. And may have been inadvertently signalling such all along. And the American rivals themselves were probably colluding as well to keep China in check.
I suspect this "history" of Deep Learning was quite dry to begin with. As it mostly featured academics. Some cloak and dagger style suspense was peppered in to make it a more exciting read ;)
He’s inherently a part of the ongoing NYT vs Silicon Valley spat, so probably being circulated more in tech circles as that plays out.