Cynthia Rudin wins the 2021 AAAI Squirrel AI Award
pratt.duke.edu
pratt.duke.edu
My fondest memory of Cynthia, however, has nothing to do with science, and everything to do with just being a kind person. We were at the NEC Research Institute's company picnic where they had an inflatable dragon for the kids to jump around within its interior. Me, Cynthia, and my wife went inside without any kids and jumped around like idiots for a while. Cynthia and my wife got bored, so I stayed behind for One More Big Bounce. With the epic bounce, I also succeeded in cracking a vertebra, nearly passing out on the spot from the pain. Eventually, I would crawl out, an ambulance was called, and I was brought to the Princeton ER.
I would have a full recovery, but I was in the ER for several hours that night. Cynthia came with us to the ER, and when she saw how uncomfortable I was on the gurney, she went back to her dorm to retrieve her favorite blanket, so that I would have even a small comfort. I am not sure how long she stayed, but I know that she was there with me longer than anyone else except my wife.
Anyhow, she's a lovely human being and I am honored and proud to have known her and witnessed the origins of her career.
Her class became so popular within the add/drop period that Duke added a second section and also doubled the attendance for each section. I'm pretty sure she went from being supposed to teach about 70 students to teaching 300. Nevertheless, her teaching was top notch, and I learned more there than pretty much any other CS class, and still rely on this knowledge today!
I too am really glad she won this award.
But Rudin's Premise is philosophical more than practical. If the problem at hand is better solved using a black box (in terms of accuracy, precision, robustness, etc), her premise says simply, don't do it. Unfortunately in the cutthroat world of capitalism, that strategy can't compete with the cutting edge.
Where Rudin's Premise is more suitable is in writing regulations to address AI app problems where social unfairness is unchecked (like the COMPAS app that advises legal authorities on meting out parole decisions without explaining its reasoning). There are many such (ab)uses for AI today in social services or policing which merit rethinking since AI-based injustice so offer bedevils the proprietary lack of transparency in such apps.
Another excellent discussion of problems like these is Cathy O'Neil's book "Weapons of Math Destruction". Too bad she couldn't share the Squirrel prize. https://www.amazon.com/Weapons-Math-Destruction-Increases-In...
It's been a while since I read her work, but IIRC one of the positions she argues for, which I find plausible, is that interpretable models can be performance competitive. For example, it could be that the only reason black box methods outperform is because they've been more heavily researched, and if we were to put more research into interpretable methods, we could achieve parity. I also mentioned a few reasons why we might expect interpretable models to perform better a priori in this comment https://news.ycombinator.com/item?id=28838321
Until that can be done, I think few outside academia will invest time or money in alternative non-DNN methods in the hope of competing with today's even superior DNN variants. There's a decade of evidence now that DNNs are incontestable discriminators in numerous domains, relative to pre-2012 ML technology anyway.
Do we know that this is due to inherent superiority of DNNs, vs just experiencing a virtuous cycle of success leading to increasing investment leading to more success?
Here is the Rashomon set argument: Consider that the data permit a large set of reasonably accurate predictive models to exist. Because this set of accurate models is large, it often contains at least one model that is interpretable. This model is thus both interpretable and accurate.
Unpacking this argument slightly, for a given data set, we define the Rashomon set as the set of reasonably accurate predictive models (say within a given accuracy from the best model accuracy of boosted decision trees). Because the data are finite, the data could admit many close-to-optimal models that predict differently from each other: a large Rashomon set. I suspect this happens often in practice because sometimes many different machine learning algorithms perform similarly on the same dataset, despite having different functional forms (e.g., random forests, neural networks, support vector machines). As long as the Rashomon set contains a large enough set of models with diverse predictions, it probably contains functions that can be approximated well by simpler functions, and so the Rashomon set can also contain these simpler functions. Said another way, uncertainty arising from the data leads to a Rashomon set, a larger Rashomon set probably contains interpretable models, thus interpretable accurate models often exist.
If this theory holds, we should expect to see interpretable models exist across domains. These interpretable models may be hard to find through optimization, but at least there is a reason we might expect that such models
exist."Rather than trying to create models that are inherently interpretable, there has been a recent explosion of work on “Explainable ML,” where a second (posthoc) model is created to explain the first black box model. This is problematic. Explanations are often not reliable, and can be misleading, as we discuss below. If we instead use models that are inherently interpretable, they provide their own explanations, which are faithful to what the model actually computes"
I am more familiar with her older work on ranking and boosting. I do not have any technical commentary to add, just a personal anecdote that she is one of the nicest, warmest person that I have met. I wish her well with utmost sincerity.
That said, no doubt that explainable ML/AI is important.
If I awarded $1 Million to a random person every year, receiving that award wouldn't make that person more accomplished, and the award wouldn't be mentioned outside the local newspaper. On the other hand an award that gives no money but consistently awards the best researcher in the field can be very noteworthy.
What the $1 million does accomplish is make people pay attention, so everyone will much more quickly reach a verdict whether this is a price worth paying attention to. But two years is a bit quick for that verdict.
Definitely!
It's a very new prize (this is only the second winner), so it's still too early to tell. But is backed by a reasonably respectable organization.
Respectability of the prize will arise mostly from the people who receive the prizes, not from the organization itself. Would you like to receive the same prize that got all these other geniuses?
The Nobel is respectable because so many great scientists got it. The composition of the Nobel committee is irrelevant, as long as they keep giving the prize to the best.
So lets say its been running for 2 years.
And the only other comparable scientific awards of such monetary value are Turing and Nobel?
Wow very generous people.
"The machine learning model is a two-layer additive risk model, which resembles a two-layer neural network, but is decomposable into subscales. In this model, each node in the first (hidden) layer represents a meaningful subscale model, and all of the nonlinearities are transparent. Our online visualization tool allows exploration of this model, showing precisely how it came to its conclusion. We provide three types of explanations that are simpler than, but consistent with, the global model: case-based reasoning explanations that use neighboring past cases, a set of features that were the most important for the model’s prediction, and summary-explanations that provide a customized sparse explanation for any particular lending decision made by the model."
I was curious about the customized sparse explanation. It looks like there is an illustrative example from later in the paper:
"For all 700 (7.1%) people where:
• ExternalRiskEstimate ≤ 63 , and
• NetFractionRevolvingBurden ≥ 73,
the global model predicts a high risk of default."
"A rule returned by OptConsistentRule is globally-consistent, in the sense that there exists no previous case that satisfies the conditions in the rule but is predicted differently, by the global model, from what is stated in the rule. In contrast, explanations (from other methods) that are not consistent may hold for one customer but not for another, which could eventually jeopardize trust (e.g., “That other person also satisfied the rule but he wasn’t denied a loan, like I was!”)"
You can see the online visualization tool her team built here: http://dukedatasciencefico.cs.duke.edu/models/
In retrospect, it's not all that surprising to me that a model such as this is able to outperform a black box like a neural network. For example, one of the things this model does which black box models don't do is enforce "monotonicity constraints" which ensure that as risk factors increase, the estimated risk should also increase. It makes sense that this would be a useful inductive bias which improves generalization performance -- if a black box model found that an increase in risk factors decreased estimated risk, it seems likely that this would be a result of overfitting (or multicollinearity gone haywire).
Of course another reason to expect simple/interpretable models to generalize better is Occam's Razor.
My big question about this sort of approach would be whether it's able to extend to the sort of unstructured data problems that deep learning has done really well on. It looks like some of her recent papers on Google Scholar address this: https://scholar.google.com/citations?hl=en&user=mezKJyoAAAAJ... (specifically thinking of the BacHMMachine paper and the Interpretable Mammographic Image Classification paper). Maybe someone else can summarize them.
"While many scholars in the developing field of machine learning were focused on improving algorithms, Rudin instead wanted to use AI’s power to help society."
"In the past, the interesting questions were around what algorithm is best for doing this optimization. Now that we have a great set of algorithms and tools, the more pressing questions are human-centered: Exactly what do you want to optimize? Whose interests are you serving? Are you being fair to everyone? Is anyone being left out? Is the data you collected inclusive, or is it biased?"
Most engineers here get like £60k salary (£3600 a month after PAYE tax), while companies they work for make billions out of their work. Not only that, but they also don't contribute back into the local communities, because they use aggressive tax avoidance strategies. Corporations need to start sharing their profits with the workers and pay taxes otherwise it will eventually spark another revolution.
What’s wrong is that these companies make billions, most of it because they happened to get to the top of the food chain.
They do.
Not that I personally believe they deserve it. For the exact same reasons I don't think a CEO is not "worth" thousands of engineers, I don't think that just because you happened to graduate in ML you are worth tens/hundred times more that the others.
Or more accurately maybe they are worth that much, but the general population is severely underpaid.
When Hinton, Krishevsky, and Sutskever sold DNNresearch (incorporated only days before) to Google, their $44 million crossed that line too, since the company had no products or IP that was independent of UToronto, AFAIK. The three were effectively "hired" as indep contributors.