https://cdn.aaai.org/IAAI/2004/IAAI04-019.pdf
It has 490 citations.
DARPA has a whole program named after it: https://www.darpa.mil/research/programs/explainable-artifici...
The real question is whether we can get some insight as to how exactly it's able to do this. For convolution neural networks it turns out that you can isolate and study the behavior of individual circuits and try to understand what "traditional image processing" function they perform, and that gives some decent intuition: https://distill.pub/2020/circuits/ - CNNs become less mysterious when you break them down as being decomposed into "edge detectors, curve detectors, shape classifiers, etc."
For LLMs it's a bit harder, but anthropic did some research in this vein.