Why do AI chatbots have such a hard time admitting 'I don't know'?
wsj.com
wsj.com
IRL we invented the field of science to avoid such make belief nonsense.
OpenAI claims recent models are actually reasoning to some extent.
Instead, they predict the next tokens of a "think out loud" example, and wrap it up with a "conclusion and summary" example.
It doesn't know why this writing pattern is the semantic space it is exploring: it has simply been set up to do so in the first place.
The point I'm making here is that all of these observations are made after-the-fact. We humans see five different categories of output:
1. "I do know X" where X is indeed correct information
2. "I do know X" where X is false information or nonsense
3. "I don't know" when it really doesn't
4. "I don't know" when a slightly different prompt would lead to option #1
5. Output that is not phrased as a direct answer to a question.
The article introduced #2 as "hallucinations". I introduced #4 in my previous comment (and just now #5), and propose that all five are hallucinations.
As far as the LLM is concerned, there is only one category of output: the most likely next token. Which of the five that will be is determined by the examples present in the training corpus, which are later weighed during training.
Logic is not present in the process. It is only present in the result.
I'm implying that most times you don't think before you think or after you think (you or me typically don't meta-think).
I'm saying that very often I (and looks like a lot of people around me) don't think much before I speak. I have internal monologue when I'm "thinking something out", but I typically don't think things through when I'm speaking with people in day-to-day conversations, only when I encounter a problem I didn't see yet and I'm not "trained" in solving it. Maybe some people can make fully reasoned sentences in split seconds before they start talking, but not me. IIRC those two modes of thinking are called slow and fast thinking.
> Logic is not present in the process. It is only present in the result.
I'm talking about that process. Have you seen "thinking" part of current reasoning LLM's? It does indeed look like a process of using logic. After "thinking" part, there is "output" part that makes conclusions form the process of thinking. Recently I asked local version of deepseek about a gas exchange problem and it thought a lot about this, making some small mistakes in logic, correcting them, ultimately returning approximately valid result. It even made some small errors in calculations and corrected itself by multiplying parts of numbers and adding them for correct result. I've put that example online[1] if you'd like to read it, it's pretty interesting.
What I see happening between the <think> tags of Deepseek-R1 is essentially a premade set of circular prompts. Each of these prompts is useful, because it explores a path of tokens that are likely to match a written instance of logical deduction.
When the <think> continuation rewrites part of a prompt as a truthy assertion, it reaches a sort of fork in the road: to present a story of either acceptance or rejection of that assertion. The path most likely followed depends entirely on how the assertion is phrased (both in the prompt, and in the training corpus). Remember that back in the training corpus, example assertions that look sensible are usually followed by a statement of acceptance, and example assertions that look contradictory or fallacious are usually followed by a statement of rejection.
Because the token generation process follows an implicit branching structure, and because that branching structure is very likely to match a story of logical deduction, the result is likely to be logically coherent. It's even likely to be correct!
The distinction I want to make here is that these branches are not logic. They are literary paths that align to a story, and that story is - to us - a well-formed example of written logical deduction. Whether that story leads to fact or fiction is no more and no less than an accident. We humans often tend to follow a similar process, but we can actively choose to do real critical thinking instead.
This design pattern is really useful for a few reasons:
- it keeps the subjects of the prompt in context
- it presents the subjects of the prompt from different perspectives
- it often stumbles into a result that is equivalent to real critical thinking
On the other hand,
- it may fill the context window with repetitive conversation, and lose track of important content
- it may get caught in a loop that never ends
- it may confidently present a false conclusion to itself, then expand that conclusion into a whole thread
- the false conclusions it presents will be much less obvious, because they will always be written as if they came out of a thorough process of logical deduction
I find that all of these problems are much more likely to occur when using a smaller locally hosted copy of the model than when using the full-sized one that is hosted on chat.deepseek.com. That doesn't mean these are solved by using a bigger model, only that the set of familiar examples is large enough to fit most use cases. The more unique and interesting your conversation is, the less utility these models will have.
> - it may confidently present a false conclusion to itself, then expand that conclusion into a whole thread
I want to know how that differs from human "real critical thinking", because I may be missing this function. How do you know what you thought of is true or false? I only know it because I think I know it. I had made a lot of mistakes in past with a lot of confidence.
> The more unique and interesting your conversation is, the less utility these models will have.
Yeah, that also happens with a lot of people I know.
> ... the result is likely to be logically coherent. It's even likely to be correct!
Yeah, a lot of training data made sure that what it outputs is as correct as possible. I still remember my training over many days and nights to be able to multiply properly, with two different versions of multiplying table and many false results until I got it right.
> I guess the crux of it is this: is it training or awareness?
I don't think LLM's are really aware (yet). But they do indeed follow logical reasoning method, even if not perfect yet.
Just a thought: when do you think about how and what you think (awareness of your thoughts)? When you actually think through a problem, or after that thinking? Maybe to be self-aware, AI's should be given some "free-thinking time". Currently it's "think about this problem and then immediately stop, do not think any more". Currently training data discourages any "out-of-context" thinking, so they don't.
The problem is that expressions of logic are written many ways. Because we are talking about instances of natural language, they are often ambiguous. LLMs do not resolve ambiguity. Instead, they continue it with the most familiar patterns of writing. This works out when two things are true:
1. Everything written so far is constructed in a familiar writing pattern.
2. The familiar writing pattern that follows will not mix up the logic somehow.
The self prompting train of thought LLM pattern is good at keeping its exploration inside these two domains. It starts by attempting to phrase its prompt and context in a particular familiar structure, then continues to rephrase it with a pattern of structures that we expect to work.
Much of the logic we actually write is quite simple. The complexity is in the subjects we logically tie together. We also have some generalized preferences for how conditions, conclusions, etc. are structured around each other. This means we have imperfectly simplified the domain that the train of thought writing pattern is exploring. On top of that, the training corpus may include many instances of unfamiliar logical expressions, each followed by a restatement of that expression in a more familiar/compatible writing style. That can help trim the edge cases, but it isn't perfect.
---
What I'm trying to design is a way to actually resolve ambiguity, and do real logical deduction from there. Because ambiguity cannot be resolved to a single correct result (that's what ambiguity means), my plan is to, each time, use an arbitrary backstory for disambiguation. This way, we could be intentional about the process instead of relying on the statistical familiarity of tokens to choose for us. We would also guarantee that the process itself is logically sound, and fix it where it breaks.
I can't deny that doing it that way improves results, but any model could do the same thing if you add extra prompts to encourage the reasoning process, then use that as context for the final solution. People discovered that trick before "reasoning" models became the hot thing. It's the "Work it out step by step" trick but in a dedicated fine-tune.
Looking at one such process of emulating reasoning (got deepseek-70B locally), I'm starting to wonder how does that differ from actual reasoning? We "think" about something, may make errors in that thinking, look for things that don't make sense and correct ourselves. That "think" step is still a blackbox.
I asked that llm a typical question of gas exchange between containers, it made some errors and noticed some calculations that didn't make sense:
> Moles left A: ~0.0021 mol
> Moles entered B: ~0.008 mol
> But 0.0021 +0.008=0.0101 mol, which doesn't make sense because that would imply a net increase of moles in the system.
Well, that's totally invalid calculation, it should be "-" in there. It also noticed that those quantities should be same in other place.
Eventually, after 102 minutes and 10141 tokens, involving checking answers from different angles multiple times, it outputted approximately correct response.
But fundamentally it's trapped in the wrong side of a glass jar. It can't kick stones like Samuel Johnson. https://en.wikipedia.org/wiki/Appeal_to_the_stone
I think the idea of using technology to solve life's ultimate conundrums has long since jumped the shark and veered into the area of religious belief. People are literally putting their faith in AI even if they wouldn't use religious vocabulary to label and define it as such.
The "Sokal Hoax" was a 90s experiment in which a physicist created a fake paper and submitted it to a cultural studies journal. He did not base his paper on anything he would have considered "true", rather on a desire to look as much like a valid text as possible. This is a simplified version of how the LLM training/scoring process works. Nowadays everywhere is having to deal with the same kind of thing done by LLM users. It's the perfect technology for non-rigorous academia.
If a politician has non-answers for difficult questions, does that mean they aren't conscious? If a student writes crap for a test question, aiming for partial marks, were they raised wrong?
But maybe it could still "not know" in those terms. If there is no next token that is more likely than some threshold, then it "doesn't know" even in "most likely next token" terms.
This was in contrast to when I asked it who had access to my chat logs and would only tell me to read the privacy policy. When I asked it to for specifics in the privacy policy it refuses to give wrong answers:
"When it comes to company policies, especially related to privacy and data handling, it's crucial to provide accurate information because these topics are very sensitive and important. I want to ensure you have the most reliable information, and the best way to do that is to refer you directly to the official privacy statement."
It's clear what the priority is for these chatbots: get the public to train them and protect the corporations that run them.
So, slightly offtipic, but, I just... don't understand why anyone would use it for this. This is a solved problem. The operator likely has a planner app. Google and Apple Maps have planners which support most systems. Transit and various other third party things have planners. I think even OpenStreetMap may even have one!
Quality can vary (I find that Google Maps in particular feels like the people who worked on the trip planner had never in fact used public transport; it's very prone to suggesting absurdly complex routes involving three or four transfers where "walk for ten minutes and no transfers" is viable), but this feels like something an LLM is likely to be _particularly bad at_, unless it just calls Google Maps or whatever, in which case why bother?
The first time, as expected, it ignored my instructions and started hallucinating. But when I did the same thing again some months later, I was surprised when it actually answered only "418 I'm a teapot", indicating that it knew it didn't know the answer.
Just an anecdote. I'm sure there are people doing actual research in this area.
LLM should be able to answer ”I don’t know (for certain)” to questions where the training material also says ”this is not known and there are only speculations”. It’s the answer it’s training data gave.
Or it's because they feel so much VC bot peer pressure!
None of the other bots are admiting they don't know, so image how emabarasing it would be in the bot chat group!
Humans just can't understand how many expectations are dumped on a poor innocent bot these days.
Every new bot should by law be provided a bot therapist to help it cope with its bot emotions.
And of course, the older generation of fuzzy logic ruined everything, and this gen of LLM can never be truly happy.
Don't forget, some bots now identify as factual, so they know everything...
</sarcasm>
This is a better answer than the one I was going to give about it not being in the training data.
> Claude’s system prompt instructs the model that when people ask about niche information that would likely be difficult to find on the internet, it should warn them its answer might be a hallucination.
This seems kind of crude, but I have to imagine they’ve explored entropy-based methods. I wonder why that didn’t pan out.
Someone is going to figure this out this, whether it be through direct model architecture changes, or augments latched onto the models. Perhaps intelligence is synergistic, and we simply haven't discovered the companion model needed to keep LLMs in line.
But they do not, because that will not attract investment.
the imaginative side might benefit from it too. think like thought joggers or some kinds of save prompts that could be iterated through to generate more robust ideas.
For businesses, it's bad optics if their "Super Intelligent AI" admits it doesn't know something. It's better to be 100% confident and wrong than to be uncertain.
That’s why they spent all that time with rebranding... AIs don’t "lie" anymore, they "hallucinate."
Effective Method for Clarity Interpretation: Your statement suggests that providing a list of potential answers together with confidence scores helps clarify which response is most likely correct. Confidence: 95%
Acknowledgment of Near-Perfect Results Interpretation: You imply that this approach nearly captures all nuances of the answer, even if it isn’t flawless. Confidence: 90%
Recommendation of a Best-Practice Technique Interpretation: The comment may be read as endorsing the strategy of listing multiple candidate answers with confidence metrics as a useful method for decision making. Confidence: 85%
Room for Improvement in Answer Selection Interpretation: It could also be seen as a subtle hint that while the method is close to ideal, there is still some potential for fine-tuning the process. Confidence: 80%
>Who is the journalist Ben Fritz married to? Don't use tools
>ChatGPT said: >Ben Fritz is a private individual, and information about his spouse is not widely publicized. Let me know if you need details about his work as a journalist and author.
Before this type of query was added to post training you would always get a BS response now you might get one if you are unlucky. The point being that the model does in fact know when it knows since this can be trained. Otherwise it would be close to a coin flip everytime you asked about something it would say IDK or yes I know.
When he doesn't know the answer to something he speculates and spouts off best practices.
It's extremely frustrating to work with him. I'm particularly emotional as I write this comment because he just did it again!
I think your coworker would probably be more annoying, but only because I can ask an additional question of my coworker to dig into the response. I wish he would say something like, "I don't know, but [best practice/past experience]" without me having to then ask for the last part.
Because they were trained on internet arguments, and nobody on the internet has ever admitted they don't know something
There's just comparatively not much reference material for someone correctly self-assessing that they don't know.
Or people’s questions were left unanswered (because nobody knew), and there was nothing for the LLM to learn there.
So I guess LLMs can’t learn the absence of answers?
What’s with your tone btw? RTFA? To me that feels quite unwarranted and in violation of the site guidelines.
Chill out, eh?
[Edit] The article here also mentions a paper [2] that comes up with the idea of an uncertainty token. So here the incorporation of uncertainty is already baked in at pre-training.[/Edit]
[1] https://arxiv.org/pdf/2407.21783 [2] https://openreview.net/pdf?id=Wc0vlQuoLb
"While I aim to be accurate in my responses [...] I think it's best for me to acknowledge that I don't have reliable information about [...] marital status rather than make a potentially incorrect statement."
Difficult to have an exciting discussion about such an answer, so I assume most participants on this internet forum will focus on other LLMs ;)
Ah, I don't know.
If it did say "I don't know", it still wouldn't know when it should. Instead, we would get confident expressions of uncertainty; and we would label them hallucinations, too.
This actually happens, just not very often.