It obviously breaks down with humans too, given we so easily hallucinate and confuse things we "know". However i still suspect we're more reliable at probing information we've experienced vs not. Even if the case of poisoned knowledge, eg a crime scene accidentally implying information to a witness that the witness doesn't actually know, we still "know" that poisoned information via incorrect inference. Ie we "experienced" it.
Wonder what architecture would allow for this style of information/weight probing for an LLM.
Isn't that precisely what the LLM training does? It signals strength of those facts, via repetition.
In that silly example/thought, the LLM would effectively need the ability to query the strength of a fact/spatial region/etc.
Right now i believe the LLM is more just the output of those weights. It has no way to inspect the strength of the signal. Eg it doesn't know if blue in "The sky is <blue>" is a strong or weak signal, it just predicted that next token.
If we could somehow encode strengths along with every token, eg "The sky is <blue:1.0>" or something we could perhaps give it a sense of [un]certainty. Though i imagine it would look differently than that since we'd want to encode this information in some sort of multi-dimensional space, rather than purely by token - since tokens aren't that valuable. Eg the "knowledge" in an LLM goes beyond tokens, and so too should signal strength.
I'm of course speculating on all of this and i have no clue on anything.
Yes, but the prediction is itself encoding that it's a strong signal.
Meaning that, if given a context (prompt), it predicted that "the sky is blue", that's because the sky being blue is a stronger signal under that context - statistically speaking.
My point was that i feel like humans have these two aspects, the ability to have a fact, and the ability to have a signal to the facts strength. I propose that as an explanation why we can internally analyze our understanding and come to a conclusion that yes, we do "know" it.
We also at times can't figure out how we "know" it, either because we've made up a detail (filling in blanks, assumptions, etc), or because we forgot where we embedded this detail. The lower the signal strength of this validation the more we can be unsure about a "fact".
I feel like LLMs are one half, but not the other. They effectively need a RAG for all of the knowledge they have, and if the counter for a given fact/idea/etc is low enough, then it's an low signal.
The question i have is how to do this efficiently. Of course i'm just speculating too, i don't know any of this.