Whereas with regards to tech, society is currently at the level of "computer says foo" (so it must be true). Requiring explanatory output to justify decisions is a likely path forward for making the use of LLM's accountable. But given how little we've been able to make traditional human-guided tech companies accountable for things like sharing scoring formulas and abusing personal information, I'm not hopeful.
In essence, dealing with plausible but potentially misleading justifications is something that we have had to do forever, and will still have to do for future artificial agents as well.
When using ChatGPT, I can't help but thinking "this sounds like a high school essay". But really, it's not. Rather, it's someone who has college+ level of reading and studying, but has never had to have their ideas tested or scrutinized. A user of ChatGPT kind of mitigates this by asking for clarification, which is a form of learning in the specific session. But in the real world, ChatGPT-as-student would be adjusted with that feedback and then incorporate it into its overall model for serving the next user.
[0] For example, the SSC post about "predictive processing" resonated strongly with me
An LLM can sometimes be prompted to give a response that looks like an explanation for how it came up with a previous response, but it is important to realize that the actual process by which it came to make the original response bears no relationship to the “explanation” it gave. Because everyone's experience with language has been with human-generated language, which is often not completely rational but, when it is not, often deviates in ways we intuitively understand, it is difficult to see the responses of current LLMs as being fundamentally different, even when we know they are.
The fact the brain reaches decisions before we are consciously aware of them does not mean the brain does not tag decisioms with motives and other information necessary to explain them, though obviously we have to trust this information, and it may sometimes be wrong, but it has a pretty good track record and there’s no equivalent facility for AI today.
This holds especially true for the kinds of actions taken by AI today which are usually deliberated thought processes for humans.
This isn’t the same thing is saying all actions are explainable, but it is quite different to being a black box. In general, one only needs to question the explanations of a person if there is a good objective reason to do so.
> It is well known with experimental evidence that when people justify their decisions with the reasoning leading to it, it does not necessarily have anything to do with the actual reasons for the decision
Out of interest, what’s the highest quality evidence you are aware of in this area?
- fundamentally limited
- rough edges sanded off in the ball tumble mill of evolution