The Geometry of Truth: Do LLM's Know True and False
saprmarks.github.io
saprmarks.github.io
GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975
Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334
Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221
>This allows us to produce 2-dimensional pictures of 5120-dimensional data. See this footnote for more details.
Read that line out loud and get a laugh out of yourself. It's something Data on TNG would say.
or dead salmon :)
They respond to prompts, scan through their data and give answers based on what their programming for probability dictates.
Humans seem to do the same thing much of the time but beneath even the most half-baked human reasoning is a self-directed sense of the world that no LLM has. The AI woo on HN is quite strong, so many will probably disagree for all kinds of shallow reasons of semantics, but even the creators of LLMs don't claim they possess anything resembling consciousness, which is necessary for understanding notions of truth and falsehood.
Are you sure?:
https://www.youtube.com/watch?v=O7O1Qa4Zb4s
> There's no reason to think it's real or required for distinguishing truth from falsehood.
There seems to be no better definition for what matter is than for what consciousness is. Hence, there is no reason to believe it's real or required to discern true from false as well?
As long as that's the case, I'd be a bit more careful with statements like these.
If we can be sure of one thing we know, it's nothing.
GPT doesn't sit there pondering consciousness and deciding, emotionally, that it thinks it's bullshit. It most certainly doesn't then decide, on its own volition, to go out into the web and comment this to others for the sake of debating them.
It doesn't sit there contemplating anything by its own volition unless its asked to. Regardless of what consciousness really is at its heart (I admit that we still don't fully know), it's a distinct self-motivated thing that we can see in our human selves and which is seen in no LLM except as a simulation produced by specific prompts.
Your first statement was that "they don't know true and false, because they don't know anything in the cognitive sense of using innate emotional and logical reasoning (faulty or not) to come to a self-directed conclusion of their own."
You haven't yet proved that assertion.
I can write a program that will print out infinite number of no repeating true statements.
> Large Language Models (LLMs) have impressive capabilities, but are also prone to outputting falsehoods. Recent work has developed techniques for inferring whether a LLM is telling the truth by training probes on the LLM's internal activations. However, this line of work is controversial, with some authors pointing out failures of these probes to generalize in basic ways, among other conceptual issues. In this work, we curate high quality datasets of true/false statements and use them to study in detail the structure of LLM representations of truth, drawing on three lines of evidence: 1. Visualizations of LLM true/false statement representations, which reveal clear linear structure. 2. Transfer experiments in which probes trained on one dataset generalize to different datasets. 3. Causal evidence obtained by surgically intervening in a LLM's forward pass, causing it to treat false statements as true and vice versa. Overall, we present evidence that language models linearly represent the truth or falsehood of factual statements. We also introduce a novel technique, mass mean probing, which generalizes better and is more causally implicated in model outputs than other probing techniques.