We ought to be able to do better than building our ML models as black boxes and saying "humans are black boxes too". We're engineering this technology, so we have an opportunity to make it inspectable. We should take it.
So humans explaining themselves is often misleading data, which may be worse than no data.
But regardless of that, while post-hoc explanations may not be that good for building predictive models of behavior, observed behavior is. We generally know how people behave, and we know the expected variance due to individual and situational circumstances. Truly unexpected behavior is rare in society, and we tend to filter it out. Truly unpredictable people get locked up and/or are given medical help, but even less unpredictable people are tested and kept out of professions and activities where that unpredictability could cause problems.
>Truly unpredictable people get locked up and/or are given medical help, but even less unpredictable people are tested and kept out of professions and activities where that unpredictability could cause problems.
Given our knowledge about e.g. Narcissists or psychopaths their behavior is not totally unpredictable. I guess a serial killer could be extremely predictable, but not something we want in society.
None of this applies to ML. ML models can fail in ways we can't easily predict, for reasons we don't expect. Viewed as minds, they run on a different architecture and on different firmware than human minds. They're alien to us. Alien like a kitten who suddenly freaks out for no reason, like ants that can get stuck walking in a loop - except more so, because we've know cats and ants for as long as humanity exists, and they're more similar to us than ML models hardware- and firmware-wise.
We can't project the same incentive structure onto software companies. They are a different scale than humans, and might be able to tolerate the possible hit to their reputation better than a human can. Their incentives are usually money-based rather than social-based.
And for the models themselves, they have no incentives, unless you count "reduce their error rate" or "be interesting enough for researchers to continue to research them". We barely know why they work. Our basis for trust is rather tenuous.