Interpretable Machine Learning
christophm.github.io
christophm.github.io
The problem is that people have particular tasks that they need to accomplish and the goal of interpretability techniques is to help them solve those tasks. This is somewhat similar thematically to how a keyboard layout can help someone achieve their typing speed goals. The two main tasks that interpretability is usually geared towards is model error detection and hypothesis generation. In model error detection, the goal is that a human wants to be able to better evaluate (improve their accuracy) at detecting whether or not a model is "incorrect" or not. Hypothesis generation is more about the generation of testable causal hypothesis that can them be manipulated to manipulate the outcome.
The real trouble in the interpretability field is that very few people actually evaluate on those end tasks. The only paper I know of that did such as evaluation had a negative result (https://arxiv.org/abs/1802.07810) in that the "interpretability" did not improve the ability of people to detect incorrect model predictions. As it stands, there is currently no reason to believe that any interpretability techniques actually work in helping people achieve the tasks they care about.
https://arxiv.org/abs/1602.04938
A simple counter-example. A really neat library called "eli5" and specifically, it's white-boxing functionality for NLP and CV applications does a great job of helping one figure out if their model is learning "incorrect" information in its prediction pipeline. The paper I link show why this does a good job of improving user trust in models, which is your second point on human-computer interaction
https://eli5.readthedocs.io/en/latest/tutorials/black-box-te...
For instance, one can run the same TextExplainer algorithm on a word embedding powered model, which encodes its features in some N dimensional space that doesn't correspond to anything that a human understands directly. TextExplainer can reverse engineer which words were most important for a classifier predicting x sample in y class, when this would otherwise be difficult to do (as is the case for basically all state of the art NLP models)
All you need to do is comb through your models erroneous predictions with Textexplainer and you can figure out what "incorrect" information is being learned. I found out this way that punctuation is pretty bad for generalizable NLP, and should be preprocessed out in most circumstances. I can take the really simple vizualizations and show them to managers and suits to justify why a model was wrong somewhere. This is value gained from interpretability
This absolutely is an interpretability technique which actually helps people achieve a task they care about. The only cavet of this method is that the white-box model which immitates the black-box model isn't 100% accurate, but its accuracy is a hyper-paramater.
PS. If you look into who the developers of eli5 are - it's an agency which takes money from DARPA and works with sigint agencies. Eli5 and text-explainer can work on any type of black-box model. Spooks want to crack open the models. That should be your signal that there's valuable information to be gained from interpretable machine learning.
For instance, the DARPA grand challenge is a impressive source of footage of robots falling over for no apparent reason...
I'm not usually a naysayer but I feel like https://youtu.be/FSubdmYGVEI?t=105
Random Forest is not explainable per say... A decision tree is explainable but an ensemble of it is not to say explainable in at least statistical models sense.
In linear models the coefficients give you a linear numerical sense and also linear associations where as Random Forest give you OOB and feature importance which isn't as clear cut. It also really dependent on the quality of data and most machine learning that aren't statistical models depend heavily on a large amount of data to overcome the model's weakness (random forest being selection bias).
Statistic, to toot my horn a little, have very vast breath and depth in inferences/explainable with huge literatures on it including different fields (e.g. econometrics with time series, biostatistic with longitudinal, survival analysis, etc..). There are also techniques on building parsimonious models too like logistic regression with purposeful selection by Dr. Lemeshow and Dr. Hosmer. And different ways to build inference models versus predictive/forecast model.
I think many people know a general idea but not in depth because I mean there's a field with many deep rabbit holes.
- I fit a complex, difficult to interpret model to a dataset, attempting for forecast my sales (structure of the dataset largely irrelevant for this example)
- I take an entry from the training set and decrease the value of some price attribute by 15%, leaving everything else unchanged
- I try to predict the sales for the entry I just created using the trained model
- What happens if the model now predicts lower sales? There is a clear relationship between price and sales volume going in the opposite direction. Would lowering my prices by 15% really lead to a decrease in my sales? How do you track what's happening in the model to create this forecast? Did I use the wrong model? Was my training data incorrect? How do you explain this to a client or to a product user?
...is called statistics. At least, truly interpretable machine learning is. The whole point of that 'machine learning' (or god forbid, 'data science') moniker, as opposed to 'computational statistics', is a way of saying "we have no idea what we're doing!" It's magic!
https://github.com/christophM/interpretable-ml-book#renderin...