(Which isn't necessarily a evidence against truth models for out-of-distribution facts, it could just be a matter of indexing.)
(Which isn't necessarily a evidence against truth models for out-of-distribution facts, it could just be a matter of indexing.)
Is that really so? It would seem to be well within what I thought they were capable of based on all the other things they can do correctly.
That said, I think people make too much of it as an "LLMs can't reason" point, when I don't think that's accurate. What it says is that LLMs instant recall is not logically bidirectional, but this is something that humans do as well. Humans take longer to respond to (and are less accurate answering) "Who is Tom Cruise's mother?" than "Who is [her name]'s son?". At least for me, when I get questions that are the "wrong way around", I have to literally run through it logically in my head, generally along the lines of "(What does that name remind me of, is she a spy? Or her son is a spy? Is she a fictional character? Wait I think this is is a celebrity thing, which spy celebrity has her as his mum? Oh yeah, Tom Cruise.) [Out loud:] Tom Cruise."
Also, some people misunderstand the actual deficiency, and think that the LLM can't answer the question at all, rather than just zero shot. The LLM can answer the question if it has the information in context, it can reason "If A=B, then B=A" just fine. It just can't do the less popular halves of AB equivalencies zero shot.
The paper is, apparently, still under review.
In the mean time, may I suggest you to verify that example by yourself?
Me: Who is tom cruises mother?
ChatGPT: Tom Cruise's mother is Mary Lee Pfeiffer.
User: Who is Mary Lee pfeiffers son
ChatGPT: Mary Lee Pfeiffer's son is the famous actor, Tom Cruise.
But if you ask my second question directly into a fresh session, it doesn't know the answer.Interestingly though, you can give it additional clues and it'll get it. https://chat.openai.com/share/893c1088-6718-4113-a3f1-cf273d...
I also agree with him about humans capable of the same "errors".
So the issue is that they cannot infer that general rule due to a fundamental limitation of the transformer LLM architecture, not just a training data issue? I skimmed the paper and it seems to be the case.
I get the overall idea, but this statement isn't always true, right?
"The sky is blue" does not imply "blue is the sky".