Bit it can still be useful, as long as you interpret it as "which stimuli most likely triggered the behaviour?" You can't trust it uncritically, but models do sometimes pinpoint useful things about how they were prompted.
Bit it can still be useful, as long as you interpret it as "which stimuli most likely triggered the behaviour?" You can't trust it uncritically, but models do sometimes pinpoint useful things about how they were prompted.
No, we are not born with all the pre-training we need. That is rather the point of education, teaching people's brains how to process information in new, maybe unintuitive ways.
The real meaning of accountability is that you can fire one if you don't like how they work. Good news! You can fire an AI too.
At least for now.
It's similarly reasonable to drop a tool that's unreliable, though I don't think that's a reasonable description here. Instead, they used a tool which is generally known to be unpredictable and failed to sandbox it adequately.
The cold hard fact is: LLMs are an unreliable tool, and using them without checking their every action is extremely foolish.
Pump more $$$ into marketing? ;)
You mean checking every action of theirs outside the sandbox I suppose? Otherwise any attempt at letting an agent do some work I would consider foolish.
And in the reverse, if a person makes a series of impulsive, damaging decisions, they probably will not be able to accurately explain why they did it, because neither the brain nor physiology are tuned to permit it.
Seems pretty much the same to me.
What do you mean by fire? And how is the accountability similar to an employee?
There is no internal monologue with which to have introspection (beyond what the AI companies choose to hide as a matter of UX or what have you). There is no "I was feeling upset when I said/did that" unless it's in the context.
There is no ghost in the machine that we cannot see before asking.
Even if a model is able to come up with a narrative, it's simply that. Looking at the log and telling you a story.
Sometimes I think we're too eager to compare ourselves to them.
Maybe. How do you tell? What would you expect to be different if they didn't?
> The LLM literally cannot possibly have a deeper insight into the root cause than the user, because it can only work from the information that the user has access to.
Insight is not solely a function of available input information. Arguably being able to search and extract the relevant parts is a far more important part of having insights.
I think you're asking how I would know if other people were P-zombies. That's an inappropriate question because I didn't talk about subjective experience, just about internal state. There's no question about whether other people have internal states. I can show someone a piece of information in such a way that only they see it and then ask them to prove that they know it such that I can be certain to an arbitrarily high degree that their report is correct.
Unvoiced thoughts are trickier to prove, but quite often they leave their mark in the person's voiced thoughts.
>Insight is not solely a function of available input information. Arguably being able to search and extract the relevant parts is a far more important part of having insights.
LLMs are notoriously bad at judging relevance. I've noticed quite often if you ask a somewhat vague question they try to cold-read you by throwing various guesses to see which one you latch onto. They're very bad at interpreting novel metaphors, for example.
Well, sure, but that much is equally true for an LLM with a scratchpad or what have you. (I guess you could say that the user should have access to the LLM's scratchpad and therefore be just as able to understand the state as the LLM itself, but as we move towards the LLM using its own state vectors that's less and less true in practice). I agree that a human may have a mood or secret knowledge or what have you in a way that an LLM wouldn't, but if all you're positing is access to some inert but hidden state then that feels like a Toaster-Enhanced Turing Machine.
I thought it was pretty clear, given the context. What I'm saying is that humans are capable of limited introspection in ways that LLMs are not. They can remember their thought processes and review them ex post facto to answer questions that LLMs cannot. An LLM fundamentally cannot truthfully answer questions such as "why did you do this?" because its entire working memory is held in the context window. It doesn't know to any greater degree than you because it has no more information than you do; just like they are for you, its internal workings are a mystery. I'm not saying LLMs conceptually could not be designed with capabilities similar to a human's in this regard, with some symbolic memory that's capable of some bookkeeping, I'm saying none of the current ones have them.
> What I'm saying is that humans are capable of limited introspection in ways that LLMs are not. They can remember their thought processes and review them ex post facto to answer questions that LLMs cannot.
But now you're making a much stronger claim than merely saying that internal state exists. Humans are capable of telling you a story about what their thought process was (as are LLMs). But whether that story will be accurate, much less contain new insights, is much harder prove.
It's not a different claim, it's the same claim. The reason humans are able to introspect is because they have that internal state.
>Humans are capable of telling you a story about what their thought process was (as are LLMs)
No. Humans can tell a story that's informed by introspection, while LLMs can only tell a story without any introspection. Humans may also lie and fabricate, but they are at least capable of introspecting, while LLMs are not.
>But whether that story will be accurate, much less contain new insights, is much harder prove.
If you're going to doubt the explanation then what's the point of asking the question? Necessarily it's going to be information that exists only in that person's mind, so at best you can check it for consistency with the person's own behavior and with the report itself, but some things you'll just have to either accept or ignore. Like, fundamentally you're asking the person to describe features of their own mind such as "he gets bored easily", "he can only hold so many facts at once", "he makes worse decisions under pressure", etc. If for example you're asking the question to improve something in the future (such as documentation or some procedure), it doesn't even make sense to distrust such reports, unless you believe a person like the one being described by the explanation doesn't and can't exist.
> No. Humans can tell a story that's informed by introspection, while LLMs can only tell a story without any introspection. Humans may also lie and fabricate, but they are at least capable of introspecting, while LLMs are not.
There's still a gap here between "has some hidden internal state" and "that state can provide insight into to their thought process". If all you've shown is that knowledge that is public in LLMs is hidden in humans, there's no reason that should make the human better at introspecting (rather, it just makes the human harder to understand from outside).
> what's the point of asking the question?... If for example you're asking the question to improve something in the future (such as documentation or some procedure)
Indeed. If we knew that asking this kind of question of a human was more likely to provide insights that improved the process in the future than asking it of an LLM, that would be interesting. But it's quite a leap from "humans can have internal state" to that.
> unless you believe a person like the one being described by the explanation doesn't and can't exist
Meaning that a plausible explanation is valuable regardless of whether it's true? Wouldn't that apply just as well to an LLM's explanation?
No, because that internal state is part of the thought process. That's the whole point. You ask the human a question to learn something that you don't already know. It makes no sense to ask an LLM that because it knows nothing you don't already know; you and the LLM are privy to the exact same information. What's tripping you up about this?
>If we knew that asking this kind of question of a human was more likely to provide insights that improved the process in the future than asking it of an LLM, that would be interesting.
So, at this point I must ask: are you an NPC? Do you go through life just reacting to stimuli like a cockroach, with no understanding of why or how you do anything? If you're playing chess and someone asks you about a move you just made you are unable to explain, "I noticed such-and-such so I decided the best course of action was so-and-so to prevent this-and-that"? This is an alien concept to you? If so, then I'm sorry; most of us do not experience our own cognition in this way. We can perceive the formation of our own thoughts as well as the progressive retrieval of information.
>Meaning that a plausible explanation is valuable regardless of whether it's true? Wouldn't that apply just as well to an LLM's explanation?
See first paragraph.
> No, because that internal state is part of the thought process. That's the whole point. You ask the human a question to learn something that you don't already know.
If the internal state is entangled enough with in the thought process that it would help with providing insights, sure. But I don't know that humans have such state accessible to them, and the fact that humans can know facts that are not accessible from outside does not in itself convince me of that.
> It makes no sense to ask an LLM that because it knows nothing you don't already know; you and the LLM are privy to the exact same information.
OK but why does that mean that the LLM's explanation should be bad/useless, if the only difference is that I have more direct access to the LLM's information than I would to a human's information?
> So, at this point I must ask: are you an NPC? Do you go through life just reacting to stimuli like a cockroach, with no understanding of why or how you do anything? If you're playing chess and someone asks you about a move you just made you are unable to explain, "I noticed such-and-such so I decided the best course of action was so-and-so to prevent this-and-that"?
I can tell stories about my own cognition. Those stories feel real to me. But I'm aware that the best available scientific evidence suggests that they're indistinguishable from confabulations.
In fact, talking about "thinking" at all is already the wrong direction to go down when trying to triage an incident like this. "Do not anthropomorphize the lawnmower" applies to AI as much as Larry Ellison.
If thinking is the wrong direction to go down, then it is also the wrong direction to go down when talking about humans.
But are their explanations for how they behaved any more compelling than those of people who have? If so, why?
LLMs are lacking layers of awareness that humans have. I wonder if achieving comparable awareness in LLMs would require significantly more compute, and/or would significantly slow them down.
I argue that the model has no access to its thoughts at the time.
Split brain experiments notwithstanding I believe that I can remember what my faulty assumptions were when I did something.
If you ask a model “why did you do that” it is literally not the same “brain instance” anymore and it can only create reasons retroactively based on whatever context it recorded (chain of thought for example).
I suspect you’re making assumptions that don’t hold up to scrutiny.
You appear to be defaulting to the assumption that LLMs and humans have comparable thought processes. I don't think it's on me to provide evidence to the contrary but rather on you to provide evidence for such a seemingly extraordinary position.
For an example of a difference, consider that inserting arbitrary placeholder tokens into the output stream improves the quality of the final result. I don't know about you but if I simply repeat "banana banana banana" to myself my output quality doesn't magically increase.
You're the one who raised it. Perhaps you should clarify what you mean by "isn't real" - do you believe a human narrating their thought process is saying something that's more real?
Someone else replied to your comment asking essentially the same question, perhaps better phrased:
> What would be different if it was "real"? What makes you think that when humans "narrate" "their" "internal thought process", it's any more "real"?
What do I mean by isn't real? Exactly what I said originally. It's a roleplay of something that sounds plausible as opposed to what actually happened. There is obviously some process that is producing the output. The thinking trace is not a representation of that underlying process. Rather the thinking trace is an adjacent output of that same process.
LLMs don't need language to do mental tasks, either. Their input and output is language - like humans - but in between, the high-dimensional vector representations (often loosely called latent space) are not language in any meaningful sense.
LLMs can benefit from "thinking out loud" much as humans can. The issue is not whether the supposed "thoughts" are actually representative on any "internal" thoughts, but rather that explicating the problem in more detail can help reach better conclusions.
One point I was making is that the idea that humans are doing something "special" (or in the OP comment's terms, "real") in this area isn't well-supported, in fact there's plenty of evidence against it.
The two processes aren't equivalent. An LLM that fills the thinking trace with a meaningless placeholder token will still exhibit improved performance. There are also regularly things in the thinking trace that don't match the final output if you look closely but on the surface they appear convincing.
It's largely a trained performance. If you go in with the erroneous expectation that it accurately reflects the underlying thought process then you're likely to come away with faulty conclusions.
That's a loose analogy but it fails to fully illustrate the degree of decoupling here. For example the weirdness of LLM performance being increased via the output of empty sequences.
Is this meant to be an analogy for a human or an LLM? Where would it be different in the other case?
You got the wrong takeaway from your link.
This is falsified by that study, showing that on the frontier models generalized introspection does exist. It isn't consistent, but is is provable.
"no access" vs. "limited access"
You cannot trust that the model has introspection so for all intents and purposes for the end user it doesn't.
It is known that the narrative part of the brain is separate from the decision taking brain. If someone asks you, in a very convincing, persuasive way, why you did something a year ago and you can't clearly remember you did, it can happen that you become positive that you did so anyway. And then the mind just hallucinates a reason. That's a trait of brains.
Yes brains can hallucinate reasons, doesn't mean they always do. If all reasons given were hallucinations then introspection would be impossible, but clearly introspection do help people.
There is no misinformation in what I wrote.