What is the closest thing to a "stack trace" to debug and understand AI neural network output? I'm new to this and curious of where to start looking to further understand how one would reverse engineer the logic or reasoning behind AI output.
What is the closest thing to a "stack trace" to debug and understand AI neural network output? I'm new to this and curious of where to start looking to further understand how one would reverse engineer the logic or reasoning behind AI output.
Transformers models [1], on the other hand, are the de-facto standard for Natural Language Processing/Understanding problems and recent studies [2] show that they can be applied to computer vision tasks. By construction, transformers let you access the so called "attention heads", which carry a good deal of information about what the model believes is important about the input. So that might be a promising route for some explainability in computer vision tasks, but we are quite a bit far away from fully understanding what is going on inside a deep learning model.
[1] https://arxiv.org/abs/1706.03762 [2] https://arxiv.org/abs/2010.11929
http://cnnlocalization.csail.mit.edu/
Can this help?
The current state of the art for analysis is ShAP: - https://github.com/slundberg/shap
ShAP is primarily an instance based explainer (one image = one explanation) but if you run it over multiple instances it is possible to gather global model insights on the data. The internals of the model are still quite unexplainable compared to decision trees or anything a human can code (horrible code aside).
There is a group at ETH doing work on adversarial attacks: - https://www.sri.inf.ethz.ch/publications/
While not directly related to explainability their work is on providing bounds on how much can corruption of the input still provide valid output. Very interesting and practically relevant as well.
Finally there is also common sense. If race was a large factor in prediction then the model will implicitly learn to predict races. I am not in medicine and do not know how much it is but if it is then the only way to not learn race prediction is to make race not correlated to the targets.