Is the reversal curse in LLMs real?
andrewmayne.com
andrewmayne.com
It seems like understanding when "is" is reversible is pretty core to the capabilities of the model, but that's different than having a lot of facts memorized. For instance, a model should be able to answer "who is the star of Mission Impossible" with "Tom Cruise" based on context or internalized-training of "Tom Cruise is the star of Mission Impossible" while not answering "Tom Cruise" to a much more general question like "who is covered in water" even if it had internalized a bit from a review like "there's a scene in Mission Impossible where Tom Cruise is covered in water..." unless it had context pointing it at that specific case.
Best to simply scale log-likelihood-based training, next-token-based training trivially contains a requirement for learning all of the subproblems that predict said, next token, and hardcoding something to get warm human fuzzies would be creating a biased estimator (and move us back towards the 90s a bit).
Models already very constantly do context-dependent token utilization, it's an autoregressive feature based on the entire stream of incoming tokens. Humans have a bias to focus on the 'last token used', this is not what language models look at.
But human language is created for and by humans. Is not then operating on language in a categorically different manner an incorrect usage/understanding of language?
For more information, please see https://people.math.harvard.edu/~ctm/home/text/others/shanno...
The version with the Weaver introduction is quite good as well, there are other versions of similar papers covering the topic from different angles, I find it to be well-worth the read.
I think anyone who's used Anki to learn anything will tell you reversal is by no means free.
고양이 on the front deck, Cat on the back.
You can see this direction and remember cat is the correct answer every single time but completely blank in the other direction. Cat on the front. What Korean word is at the back ?
Conditions I mean: 고양이 => <concept of cat>; cat => <concept of cat>; either [고양이 => cat || cat => 고양이]. Under these conditions 고양이 == cat && cat == 고양이 should be known.
Your conditions: 고양이 ~> <concept of cat>; cat => <concept of cat>; [ 고양이 -> cat && cat ~> ...고양이?] (or something, trying to express you're preferentially learned one direction, usually mapping back to mother tongue while learning)
By no means are we strictly invariant. But I'd say that a good measure of your skills are in fact to be invariant. But also like you are pointing out, we also kinda train ourselves that way because it creates stronger pathways. Training methods are important too! And augmentation is critical to both human learning and machine learning. They're specifically critical to generalization abilities. You'll also see this in humans. I notice test focused educations tend (there's no "always") to create less generalization of concepts as you're over learning a thing presented in a specific way rather than focusing on the concept as a whole. Form of metric hacking if you will.
Side note: if you have tips for an extreme beginner learning Korean I'd really appreciate them. Trying to reduce the burden on my gf having to translate everything for me when we're around other Koreans.
If you have learned Hangul, then you can start with Lingodeer(app), TTMIK(talk to me in Korean)’s Grammar books, or Billie's beginner playlist - https://www.youtube.com/watch?v=sx0yyQqkpqo&list=PLbFrQnW0BN....
These are all extremely beginner friendly. You can just look at the options and choose whichever works best for you. I personally tried lingodeer and ttmik both. I did a bit of lingodeer and then ttmik levels 1 to 3. I didn't try Billie's playlist, but he's a great teacher, so I recommend it if you think that would work better for you.
So in the beginning stages, it doesn't really matter so much, but eventually you're going to have to start thinking about how you’ll be learning the big 6, vocabulary and grammar (what I’ve spoken on so far covers grammar and TTMIK takes you pretty far in grammar if you stick with it) then speaking, listening, reading, writing, because believe me, there isn’t as much transfer learning as you would expect or hope. I mean, you don't have to stress over it or anything, just keep in mind that you have to eventually think about how you're going to address those things, and the sooner, reasonably, the better.
So this is kind of where my training diverges from most traditional approaches. I know I used Anki from my example and a lot of learners use that with pre made decks from the community for vocabulary, but honestly, I just couldn't use it then.
Part of it was because at the time, learning Korean words was like drawing blood for stone. For me, it was so hard. Even five words a day could be a challenge. My brain was like, “what the hell are you doing?”. It just wasn't going in.
Come to think of it, you know grokking in NN’s? Learning a language, you quickly realize the same phenomenon can occur in our biological counterparts. It’s like one day I was struggling to learn these new words, I don’t do anything drastic or different and suddenly the next day, learning new words felt so much easier, it’s crazy. Progress isn’t necessarily linear that’s for sure.
Part of it was also because I just didn't like the Anki experience, memorizing hundreds, thousands of words in isolation. It just wasn't really doing much for me. So, I just kind of dropped it. If traditional anki works for you then go for it but gin, you’ll need to figure out the remining 4. No matter how much isolated grammar and vocabulary you learn, you’ll need to practice doing each of those things for it to sink in.
The typical process is to learn a few thousand words and up to intermediate grammar before sinking into the 4. Makes things a lot less frustrating.
What I found was this amazing audiobook series. It's a graded reader series. 100 books with audio, all free(start from the oldest). https://audioclip.naver.com/channels/57.
Here the books are further divided into stories. At first, each story is really a few pages but it does get longer and longer as the series progresses. graded readers are specifically simplified literature created to help language learners progress.
It's also so that you can read in your target language much earlier than you otherwise would have because, well, you might see a lot of people say, “oh, just read, just watch or read children's shows or children's books.” That's a trap because like, children are native speakers man. Don't make the mistake of thinking that they aren't.
The complexity of the stuff they read or watch or whatever, is well beyond the beginner learner. So if you go into children's media and learning some few hundred words thinking, “oh, this should be easy, then, yeah, it may become very demoralizing because it won’t be in fact easy”.
And just to clarify, I mean children’s stuff not baby stuff. Yeah, baby stuff is simple, but you're going to get bored of that very, very fast.
So, yeah, it's a graded reader series and the idea is just to get started. Simplified stories you might find more interesting than baby stuff.
So, I really like reading. And I don't know, I just felt like it would be better because, you see, what Anki does is SRS, spaced repetition system, right? But when you think about reading, it's really a natural form of that. So, I just felt like, okay, this is what I'm going to try.
So, I was even thinking, oh, I could use this to knock almost all the birds with one stone, you know, it’s vocab, grammar, reading, and listening practice at the same time.
But the book series, it starts with a pretty high level. I don't know if you're familiar with the CEFR levels. It starts at about a low B1(bottom end of low intermediate) or so and ends at a high C1(top end of low advanced).
So, yeah, coming from where I was, where I knew at best, maybe 100 or 200 words and jumping in, it was difficult. And honestly, I would not have been able to do it without Mirinae. So, Mirinae is a grammar parser - https://mirinae.io/.
And what it does is, so basically, Korean and English are so far apart that you could know the meaning of every word in a sentence, and you would be absolutely lost as to the meaning of the sentence itself. So, what Mirinae did, or basically took what would have been impossible for me and made it achievable. I learnt a lot of grammar this way. You place the sentence in, and it would basically parse the grammar on the sentence.
You could see the explanations of different grammar parts, different nouns, objects, stuff like that, how this affected this and all that.
It was very, very, incredibly helpful. To to this day, I think that at least for distant languages, a good grammar parser is by far the most useful tool, better than a dictionary, better than just a translator.
So, the first story, I was pasting every sentence. So, it took me about a week to get through it, a few hours every day. The amazing thing was that you could the progress, the reduction in time it takes you to complete things, because the second story took only a few days in comparison. By the third or fourth story, I think it took only a few hours. I mean, I didn't finish it in a day, but I could have if I wasn’t a bit lazy.
So, yeah, that's the benefit of graded readers. Whether you start this as early as I did or not, it’s genuinely an amazing resource. Earlier this year, I came across this Anki deck series with natural audio -https://ankiweb.net/shared/by-author/374470252
What the guy managed to do was, he took a machine learning Korean speech dataset, and he kind of configured a lot of the sentences, which were, you know, a single woman speaking sentences, about 12,000 or so of those sentences. So, he reduced it to about 7,000, and he rearranged, as much as possible to be i+1 comprehensible. i+1 is just basically the idea that you come across a sentence, and there’s only one word you don't know.
So, it's much easier to grasp the meaning of that sentence quickly and easily. Then the next sentence, only one word, then the next sentence, only one word, then, and so on and so on. Now, this deck isn't perfectly i+1. But going through it, it's very good. So, you know, you could also explore that too after the beginner grammar resources, I suggested.
I’ve also noticed these guys, https://umiapp.co/.ly. It’s an app tailored for learning words in context, comprehensible input style. I really liked the looks of what I saw from the other languages. Korean isn’t there yet but is coming soon. By the time you’re ready for it, it’s hopefully there.
Anyway, best of luck.
But are they? Unless I'm misunderstanding what you mean here, I'd say the opposite is clearly the case - learning "A is B" doesn't automatically mean learning the reversal; you either have to learn it from external source, or infer and memorize it - both involve an extra cognitive effort. This maps to LLMs as training, or "in-context learning" and then adding the inference to the training set.
In the non-simplified case I'd actually say that we have some natural experiments, but that might spoil this as well because you can in fact say that this is training for such invariance (which I'm unsure is problematic to my claim but pushback is definitely welcome). If we're learning history we'd often learn information such as "On December 7th 1941 Japan attacked Pearl Harbor and this was a major event that got America directly involved in WW2." But then on the test we'd see questions such as "What important event happened December 7th 1941?", "What was the date of the Pearl Harbor attack?", "What major event led to the US becoming a direct player in the war and when did this event occur?" and so on. Granted, because we do such testing I'd expect GPT to be able to reasonably answer many of these same questions but I'm very much under the impression that humans are far more robust/generalized and would quickly learn these reversals. And to stress, conditioned on knowing the forward direction to a strong degree. If you're fuzzy about the forward direction I'm not surprised if the reversal is even more fuzzy. Definitely biased. But more importantly, I do not think we have the same strong algorithmic bias in our encoding methods that transformers have. That's my main claim, just to clarify.
But I'm conjecturing a lot here. I think these are reasonable and we have some strong evidence for this. Maybe this is more situational than I'm presuming and I'm definitely working off of biased sampling. So discussion helps refine the ideas.
Similar criticisms are "well there are many descriptions of game x it would have read so why doesn't it play x well". But gradient descent is a dumb optimizer. Training itself is not actually like someone reading a text anymore than evolution is like someone thinking about the best way to augment an organism.
a "smart" optimizer would look at a reversal applicable sentence and know exactly what bunch of weights to change to store it in such a way as to be recalled reversibly in the future.
Inference may be smart (GPT-4 can reverse in context just fine, potentially play games from only a description) but the training is not.
> You could potentially describe how a game is played to GPT-4 without examples and get it playing it correctly but passing that same description into the training process of the model just gets you a model that can describe your game correctly.
It seems like you're referring to an associative Hebbian/Hopfield-like lookup, which the current 'dumb' optimizers already do. Better yet, the learning rates are normalized by the diagonal of the empirical Fisher so that the learning w.r.t. to some estimated expected information is more constant, for said associative lookup operation.
Additionally, the training loop (which you call 'dumb') is just...a teacher-forced version of inference, which you call 'smart'?
It's better to simply minimize the log-likelihood in a scalable way. Hand-engineered solutions rarely survive compared to strong-scaling ones.
>Additionally, the training loop (which you call 'dumb') is just...a teacher-forced version of inference, which you call 'smart'?
What powers In context learning is not very well understood but it doesn't appear to be or really work exactly like just a non teacher-forced version of training. There are qualitative differences. The same models have no problem with this 'curse' when the information is provided in context for example.
>It's better to simply minimize the log-likelihood in a scalable way. Hand-engineered solutions rarely survive compared to strong-scaling ones.
I never said anything about it being bad.
I mean, yes. One distills information from a training set into a compressed representation, and the other generates a compressed representation that yields (more or less) fixed state space attractors. It's just inducing a bias over the state space of the network, nothing incredibly special, though I'd consider the initial stage of training to be the most important, as it is responsible for all of the ingest of all of the embedded information the network will be using during autoregressive inference (especially w.r.t. the context of doing so during longer generations).
So the notion of 'in context learning' is one I find to be a bit of an illusion, of course, as no actual learning is being done, just the induction of biases, which appears to give rise to a transiently-'better trained' network.
You could see this as a bit straightforward perhaps, but I feel it needs to be said.
(Perhaps though it has to be done during training so that "Who was the ninth Chancellor of Germany" becomes knowable in the first place.)
No, in English this would be “apples are red”, not “the apple is red”.
"An apple is red" does not imply that "Red is an apple". Use of the indefinite article makes the non-equivalence clearer.
Not to Bill Clinton this, but I think you are considering this for the wrong interpretation of "is". In this case, "RMS is the founder of the FSF" is the sort of statement implied here. Nobody else was or ever will be the founder of the FSF. So yes, for what the author is describing, "B is A" should always hold. B, in this case, isn't a category, it's an exact identity. The author doesn't explicitly state this, but it's clear this is what they're talking about.
This problem exists for both LLM's and humans, I think it may be fundamental to reality itself, as it currently is at least.
I wonder if they changed the training data to consistently use "equals" where possible (even if it sounds weird to the reader) would make a difference.
I'm not going to hype the technology more than it needs, but it seems to do a pretty damn good job of just this. And to the author's point, it does get this right.
And in fact, I'd suspect many of the cases where the LLM gets this wrong are caused by lots of examples in the training data which were written at different times, where someone else was chancellor. Or where the LLM gets numbering wrong (which is a well established problem).
And this is just one thing that makes sorting out what's going on here difficult - essentially, we have biologically LLM's [thinking they are] decoding silicon LLM's, of course the results are going to be weird and counterintuitive.
[1] Examine the specific language you are using, but also the colloquial ~interpreted/virtualized assertion that one is left with, if they are reading your comment "in good faith".
They implicitly recognize (and model) patterns in text, then continue those patterns.
While the patterns that an LLM has modeled may align to the high abstraction we call language, LLMs actually work at a much lower level of abstraction: plain text.
It's the content of that text that is significant. We encode language into text, and LLMs infer something "close enough" to language by blindly modeling it.
The fact that the current ones don't is surprising.
Unlike humans, LLM's can consistently realize (or at least claim to realize) when they've made a mistake (like assuming one's beliefs are necessarily correct, [because some persuasive story]).
(1) Mortal was Socrates.
(2) All humans are mortal.
(3) Therefor all humans are Socrates.
Or in short: LLMs will handle "The apple is red. Red is the " just fine, because the first part is heavily biasing the answer. They won't often complete "Red is the " alone with "apple", because why should they?
The entropy gotta get slung around somehow.
>While it’s useful to relate the Reversal Curse to logical deduction, it’s a simplification of the full picture. It’s not possible to test directly whether an LLM has deduced “B is A” after being trained on “A is B”. LLMs are trained to predict what humans would write and not what is true (Lin et al., 2022). So even if an LLM had inferred “B is A”, it might not “tell us” when prompted. Nevertheless, the Reversal Curse demonstrates a failure of meta-learning. Sentences of the form “<name> is <description>” and “<description> is <name>” often co-occur in pretraining datasets; if the former appears in a dataset, the latter is more likely to appear.4 This is because humans often vary the order of elements in a sentence or paragraph.5 Thus, a good meta-learner would increase the probability of an instance of “<description> is <name>” after being trained on “<name> is <description>” . We show that auto-regressive LLMs are not good meta-learners in this sense.
Correction. The Model can. Training cannot.
It's not that humans aren't suffering from this too. It's possible to spot it when you're learning, or even easier, when you're teaching someone a new thing. A person who learned that "A is B" does not automatically learn that "B is A"; they need to first process it, perhaps run the reversion explicitly in their head. I'd say that checking if the student can infer "B is A" having learned "A is B" is a good way to tell whether they're starting to comprehend the material, vs. just memorizing it.
This isn't surprising if you think about how the machine actually works rather than treating it like a sacred magic box.
The default assumption for why there is any "success" for in-context learning should be that it's just picking a nearby token that "fits" not a process of logical deduction.
edit: The obvious band-aid-fix is to feed reversed sentences into the training data without telling anyone, after which LLM boosters will proclaim that LLMs "learned" to reverse logical implications.
I don't know in what way this is some trivial "recall" problem as you suggest. Does the reversed implication in question need to be in the training data explicitly or not? If it does I don't understand how you can claim a logical inference capability.
No that's not what the paper says. It says training won't immediately store this information in a way as to make it reversible.
Let's get one thing straight. This is a common problem for human learning as well. Anyone who has used Anki for language learning will tell you that if you just train on target language word on the front and native language word on the back, you will fail the reverse unless you specifically train for that.
This is specifically a problem of recall. Not a problem of making the logical inference. If you ask the human language learner immediately he has learnt the new word for the reverse direction, he will obviously tell you the correct word.
But later he may not recall the reverse even if he remembers the original. Again this is a problem of recall rather than the ability to make logical inferences.
In the same vein, if you give the LLM the original direction in context and immediately ask the reverse, it will correctly tell you. It can make the logical inference.
And although it's a complete waste of time because the answer is obvious when you do do these types of adversarial experiments, like adversarial blockworld, you find that LLMs are not internalizing abstract principles inherent to reasoning and are instead doing something like string-pattern-patching.
If GPT4 rarely makes those errors it's clearly not an insurmountable problem at this level of training, but it does imply a lot more difficulty fine-tuning models to reliably retrieve correct information from specific text.
I don't know that I even want these things to generalize. I get it in academic terms, but from a CTO perspective how much this matters to the business? We already know our policies and procedures and aren't exactly excited about the idea of something getting curious with flipping things around.
I don't want it to generalize and start making weird assumptions like "some men are tall". I want it to say "I don't know", or provide a statistical signal indicating the same.
LLM's have a specific architecture that makes them good at approximate retrieval. They have a specific learning process that makes them able to do this over a vast training set.
Doing things like reasoning about facts is possible using either symbolic or specific sub symbolic structures, but combining these with an LLM is a bit tricky. It's something that could be done for a demo in a 6hr/day/week project (depends on the demo) but to do it for real in an application... or as a robust bit of science... well... 6 people for 6 mths/years/?
The opposite statement, humans don't know when "is" means equal is also a true statement, perhaps even more true.
I doubt MS would put their service in lot of their products.
I consider them the target audience for upgrades, I think the ROI could be massive.
Can you explain what you are skeptical of? There is ample evidence that LLMs are what they say they are: a series of transformer (usually) functions with learned weights that can successfully be used to generate human like text in many domains. Do you disbelieve this? Or are you skeptical of something else.
This is an interesting point, and made me think on whether this "reversal curse" is something we experience with our own, human neural networks. I think it is. Like, I can imagine being given a character in a movie, being able to tell you what actor played them, but given the actor and the movie, not being able to tell you what character they played, or vice-versa. So the pairing exists in my brain somewhere, but I only have an accessible pointer to one side. I think we run into cases like this all the time, actually.
Another funny limitation: when a word is "on the tip of your tongue", what often happened is you thought of it, but your brain rejected it, so now it's on "cooldown" before it appears in your mind again. Usually this mechanism helps but a bug in your brain causes it to be harmful. If you start thinking of something else it "refreshes the cache" and the word comes to you.. we're really not that much less buggy than the machine.
disclaimer: there's other theories about why the phenomenon happens
For example when trying to learn immunology, "This bacterium causes neural bacteriumeritis", score 100% on the set of cards after a few sessions.
Add "What are the possible causes of neural bacteriumeritis?" and sometimes I'd draw a blank.
There's a lot of nuance in how we learn facts and their relationships.
It reminds me a lot of data structures in programming to be honest. It's a lot easier to retrieve a person's age from a dictionary indexed by first-name+last-name than from a list with 10000 objects with people in them.
On another note I should have read the actual paper mentioned in the post critically instead of just skimming it. I completely missed the footnote about the reversal curse not being a problem if everything is present in the initial prompt
If you deviate a little from the examples in the article GPT-4 gets it all wrong:
Who is the eighth Federal Chancellor of the Federal Republic of Germany? - Olaf Scholz (wrong, Angela Merkel)
https://chat.openai.com/share/937795ea-bd91-43ee-bd76-1a125f...
Who was the eighth Federal Chancellor of the Federal Republic of Germany? - Helmut Kohl (wrong, Angela Merkel)
https://chat.openai.com/share/937795ea-bd91-43ee-bd76-1a125f...
Interestingly whether you put is or was into the question does make a difference.
Even when allowed to surf the web, ChatGPT gets it wrong:
Who was the eighth Federal Chancellor of the Federal Republic of Germany? - Gerhard Schröder (wrong, Angela Merkel)
https://chat.openai.com/share/867cb3bc-f642-4d2a-b80d-fe8dc0...
Though it's referring to the right Wikipedia page: https://en.m.wikipedia.org/wiki/List_of_chancellors_of_Germa...
> Who was the eighth Federal Chancellor of the Federal Republic of Germany? - Gerhard Schröder
Its probably counting Walter Scheel, who was Acting Chancellor in 1974, and you are probably not.
When I asked ChatGPT to list the chancellors in order and identify the eighth, it listed eight ending in Schröder, with Scheel and his ten day Acting Chancellor tenure in May 74 as number 5. (Scheel’s tenure is in the timeline on the Wikipedia page you cite, though not in the numbered list in that page, which is why there is a dicontinuity in thr dates on the numbered list.)
Schröder is the only answer it gave when using functionality that would bring some representation if a list into its context first, and 10it did it with different prompts and mechanisms for bringing a list into its context.
That LLMs are bad at counting-related tasks without doing that is well-known, and not a point I felt needed belaboring.
Also, knowing German chancellors was always easy, so that fact that a machine falls behind any normal dictionary or the German Chancellor's website is poor form.
It's like a calculator getting basic addition wrong. I don't want a future full of poorly performing machines.
It's a technicality, but a valid technically and it's a technical question. I agree that it's wrong in a relatively small and surprisingly human way, but it's wrong in a way that including a red apple amongst apples is not.
The machine is a complete moron for not being able to get basic data from basic sources right.
> Chancellor of Germany, Acting, 7 May 1974 – 16 May 1974
(And so does the German Chancellor: https://www.bundeskanzler.de/bk-de/kanzleramt/bundeskanzler-...)
There is zero ambiguity about who was Chancellor and who was not.
Even on https://en.wikipedia.org/wiki/Chancellor_of_Germany he's listed as "Vice Chancellor Walter Scheel served as acting Chancellor from 7 May to 16 May 1974" between 4 and 5.
There's obviously some ambiguity to it considering we are two humans discussing this with a claimed discrepancy between English Wikipedia and German Wikipedia, but your conclusion is still that the AI is spitting nonsense.
The way to be Chancellor is through article 63 of the Grundgesetz, while Scheel was put into the caretaker role via article 69. This explains it a bit https://de.wikipedia.org/wiki/Vizekanzler_(Deutschland) - Scheel was only taking on the function, not the office.
This kind of giving some machine the benefit of the doubt when in fact there is zero ambiguity is really a path that makes me think we will have mostly machines designed for marketing and other non-critical things.
When you need to bring up article 63 and 69 of the Grundgesetz to prove that the claim is ludicrous, maybe the reasonable thing to say instead is "I understand why you might think that".
I honestly don't understand why in situations with a clear ground truth there is a need to debate and why we would want machines to bungle that.
So either need to be fully descriptive (e.g., something like fulfilling the functions of the Chancellor, while not ever having the office) or it will be also open to being misunderstood.
The issue here really is that the German succession doesn't ever transfer the office, but only the function (which is different to the US, for example). So here Scheel followed Brandt, but not into the office. Only someone having the office is a Chancellor and there is a specific way to that office.
I think there’s also a functionally useful way to describe someone’s role as what it functionally is even if it’s not legally that. You’re point is super well taken, if the machine is intended to be a fact oracle, it’s awfully loose and adds a lot of interpretation in areas of ambiguity.
I would say IMO that’s specifically the power of these machines. An awful lot of human endeavor doesn’t require literalism but semantic approximation and interpretation that machines were literally incapable of. Its weird language is enough to achieve that, but I think it’s overly restrictive to assert a broad lack of utility in critical systems. An awful lot of critical systems actually need more “probabilistic” interpretation than literal fact oracling.
I'd agree that often things human are more loose but when there are narrow definitions ignoring them is dangerous.
I'm just going to point out that you're calling an AI moronic for not getting this right when a bunch of humans are also disagreeing with you. Frankly, this really undermines your case that this is an AI failure.
I think this is really an issue of semantics & translation, because that is just what 'acting X' means? (Admittedly not always while still doing something else too, but I don't see that as significant - if it had gone on long perhaps a junior minister in the foreign office would've been named acting FM in his place too.)
Article 69 Grundgesetz has a sort of "Caretaker Chancellor" that has the function but - importantly - not the office. They way to have the office is through article 63.
It works differently in Germany to the US, for example, where the vice president could become the president, while the vice chancellor only ever gets the function, not the office (unless through a proper vote for Chancellor).
When there is actually ground truth, machines should be able to recover that, not take weird turns.
I'm British, not American, so I'm not assuming something like the vice president automatically becomes president, we don't have that either. If anything it's even less than Germany since (at least in theory and history, modern media etc. makes it a bit different in practice) there's nothing special about the PM, it's just the governing party's leader. I suppose though you could say the heir apparent to the throne immediately becoming the monarch on the death of the previous one is like vice taking over - but I don't think that detracts from my point because nobody would call that 'acting monarch', they just are.
'Acting X' means doing the necessary duties, but not any major decisions that can be avoided/deferred, while the 'real' replacement is found. Sometimes it ends up being the same person, e.g. the head of some division is 'acting CEO' for a while as the board searches for a new CEO, ultimately ends up going with that person and title changes to just 'CEO' - or they don't, and go back to old job, or maybe quit in a huff, and the person they found is just 'CEO' taking over from the 'acting'.
There is no named role of "Acting Chancellor" created in the basic law, instead someone is just tasked with performing certain duties (and people might sometimes call it "acting Chancellor" or whatever). However, the position "Chancellor" is a well defined & special role (unlike perhaps PM) and whatever that other "acting" role is, it isn't a Chancellor. It is a bit like adding Oliver Cromwell to list of English Kings or Kamala Harris to the list of US Presidents - you can do it but you then move beyond the "standard" definitions.
I don't think in English usage there is any meaning attached to 'acting X' which 'caretaker X' (as you're happy to call it) doesn't also carry. Both are used interchangeably, the former you'd put on your CV, the latter might be used by the media when your employer put out the less release announcing it, but same thing: for some reason there is not currently an X, but you are fulfilling some necessary duties that that person would do in the meantime.
https://chat.openai.com/share/ddd2800e-630c-4ff7-9992-14b265...
Technically there's no right answer to this question. "Is" implies present. But Angela Merkel isn't the present chancellor. Olaf Scholz is. But he's not the eigth.
Am I missing something here? This paragraph reads like complete gibberish to me.
Also, I don't buy the experiment at the end. If you fine-tune the model with exclusively Tom Cruise data, I want to see proof that it doesn't just answer "Tom Cruise" all the time. I want to see that it says Tom Cruise wrote Aces in the Stream, but doesn't say Tom Cruise wrote Kings in the River.
Spoiler: It doesn't say "Tom Cruise."
I'm thinking this might work because "equals" is always logically reversible whereas "is" is not, assuming the statements are technically truthful.
Neural networks are graphs insofar as they're networks of connected neurons and not further(definitely not graphs of information as the writer seems to think). Given the fact that neural networks are quite literally an n-dimensional function once trained, it's more accurate to call them an "equidistant grid of points" than it is to call them a graph of information since literally all they do is take an n dimensional vector and output an n dimensional vector.
They may also be getting it mixed up with semantic search, where an entity like that would exist in a graph and have connections to related concepts.
[0] https://towardsdatascience.com/first-neural-network-for-begi...
A common argument is that LLM by their nature only model the former, not the latter.
This suggests that you may need to train an LLM on its own output. Goes against the conventional wisdom!
I don’t know French, but I am sure there are many cases where an English phrase or sentence that includes ‘cat’ should not be translated into French with ‘chat’ and vice versa. (When asked, GPT-4 offers two such examples: "let the cat out of the bag” and "avoir un chat dans la gorge.")
I don’t mean to suggest, though, that memorizing word pairs is not a good way to learn vocabulary in another language. For me, it was an essential step in acquiring the second language that I am now fluent in (Japanese).
More importantly perhaps, it's not how people (at least non AI/ML research types) typically think they work, and much of the handwaving and hype around it isn't helping improve that.
Indeed, just because the lack of that capability is a result of the design of LLMs doesn't mean it's a feature of LLMs. It could also be that it's a bug of LLMs. Which one depends on what the expected behavior is, and the expected behavior from the product is being able to perform the above logical derivation.
tl;dr: the article reaffirms that the "curse" is true, and for the reasons claimed too, but it's a feature, not a bug.
The people making the argument about reversal curses are not logicians, and most of them don't know anything more about formal logic than what anyone would pick up in an undergraduate "Intro to Proofs" course.
That said, the semantics of the word "is" in natural language really doesn't matter in this debate. The semantics is a red herring, if you will (while a red herring is not semantics).
After all, LLMs cannot learn "the quantity B is mathematically equal to A" from examples of "the quantity A is mathematically equal to B" either, even when the rest of the corpus clearly explains that this _is_ in fact always reversible.
As a logician, do you consider this necessarily factual:
- using ternary logic
and/or (I am asking both independently and in combination)
- considering the numerous possibilities for ambiguity in "doesn't matter"?
'I’ve been playing around with fine-tuning LLM models for years and still don’t have any hard and fast one-size-fits-all rules to apply. Every dataset lends itself to a specific way of training.'
But this is precisely the reason why nearly every claim about what "AI" will soon be able to do, is misguided. The failures of LLM's are exceedingly unintuitive to anyone who doesn't have a lot of experience with them (and maybe sometimes to people who do).
Spreadsheets can be used by people who don't know how spreadsheets' internals work; after a bit of training (in my experience, about ten minutes) they can get a decent intuition about how to use a spreadsheet (at least for the simple stuff). The same is true of well designed web pages, word processors, music players, etc. Even software with more complex training requirements like CAD, statistics packages, video editing, etc. will usually behave in a more-or-less intuitive fashion for a user of the appropriate target group.
LLM's fail in unexpected (and sometimes hard to spot) ways, and are being marketed as if they are a tool for the general population to use, when their training (and failure modes) are still a "dark art" even for people with years of experience in them. That is not a feature.
But of course "how neural networks function" is that they fail at basic logical deduction and do not generalize.
So again, if I'm reading it correctly, he's hand-waving away the inability to make basic logical deductions because that not something they can or should be expected to do. As I read it, that means the reversal curse only exists if the answer to the question "can LLMs do logical deduction?" is "yes". If one takes the position that LLMs can't do general logical deduction, which seems to be the author's point of view, then there's no expectation that knowing "Tom Cruise is the son of Mary Lee Pfeiffer" is sufficient to determine “Tom Cruise’s mother is Mary Lee Pfeiffer”.
Am I missing something?
These are the salient takeaways I got:
- Is/Was wording might matter. This is something probably a bug.
- 30 facts about a person might simply be too little for "B to A" generalization
- Extra precision/context in the prompt can help locate the "B to A" inference.
- How you cut up your training data can bias inferences in surprising ways.
- "B to A" generalization clearly does happen, even without "B is A" in the data, but it's not as stable as you'd want.
> it's not as stable as you'd want.
Which I take to mean that a model can sometimes confabulate "B is A", solely out of random variation, and that it's possible to bias the data and prompt to generate the expected response. The model hasn't done any logical deduction, the response is just a bias-influenced lucky break.
1. That is, the LLM should be able to learn that “A is son of B who is female” implies that “B is mother of A”, regardless of who A and B is. It should then be able to apply this pattern to A = “Tom Cruise” and B = “Mary Lee Pfeiffer” and deduce “Mary Lee Pfeiffer is the mother of Tom Cruise” without even a single example.
But the association between the words ‘Mary’, ‘Lee’, ‘Pfeiffer’ and ‘son’ or ‘mother’ point to other words more strongly than they do to ‘Tom’ or ‘Cruise’. Sure, Tom Cruise is probably in there but so is probably John Henry Kelley (Michelle Marie Pfeiffer’s son), and George Washington Custis Lee (whose father was Robert E Lee and whose mother was called Mary).
A random piece of text containing the words Mary Lee Pfeiffer is just, based on its training, not likely to be about Tom Cruise. There’s nothing there anchoring it particularly to those words.
It’s possible that the ‘Pfeiffer’ in there screws it up more, as well - it’s like it vaguely knows there’s a Hollywood connection but of course it crosses its wires and thinks it’s probably Michelle Pfeiffer.
Be honest: if a pub quiz question came up asking “which Hollywood superstar’s mother is Mary Lee Pfeiffer“, you’d probably not guess Tom Cruise either.
But here’s the thing: once we go beyond training into specific completions, if you ask GPT ‘who is Tom Cruise’s mother’, it answers correctly; and then once it has that A is B in context, you can ask it ‘who is Mary Lee Pfeiffer’s son?’ And of course it knows how to complete that B is A.
So I definitely agree with the piece here - there’s no ‘reversal curse’, there’s just asymmetric information relevance.
However, you are right that it's not always correct since "Poetry is Literature" and "Steam/Ice is H20" are both examples where this breaks down.
My annoying pedantry for the day.
Jeff is the short haired guy != the short haired guy is Jeff.
It depends on context. You can’t fully equate them.
The actual paper explores equivalences that actually make sense, and whether LLMs have trouble with them.
At least according to some people's expectations, presumably.
So I find it weird to call this "a feature" and also act like this is surprising.
But can I also take a minute to just say I really hate GPT experiments? They're performed on a stochastic model, with proprietary weights, proprietary training data, proprietary training methods, and above all is constantly changing and at a rather fast pace. It makes for a very non-scientific process as you can't decouple a lot of important factors and reproduction is a crap shoot. It is not a very good way to go about studying "how __LLMs__ work" and is rather "how does GPT work at this particular moment in time and aggregating all these unknowns?" There's some bitter sweetness because I do like that GPT is free but it feels like a major edge that they have is that the community just does a lot of free research for them and in ways where they can better interpret results than the people who performed the experiments in the first place. I really do believe that academic works shouldn't be focusing on proprietary and dynamic methods. It's fine to include them in results (and preferably added post review to avoid identity spoilage (or we openly acknowledge that double blind is a joke)) but I'd rather most researchers focusing on the general concepts and with the ability to dive down the rabbit hole than playing a wack-a-mole game.
Also, I'd totally love it if more research papers were instead blog posts. Kudos to anyone posting their research on blogs, academic or not (I don't care about your creds, your real creds are your work). Papers are about communicating with fellow scientists, right? Why do we need journals and conferences these days?
Good writers, marketers and manipulators know this about people. You seed the ground with the concepts you'll need to introduce, so they're present in the person's 'model' and you can elicit them later, on request.
Formalizing things in the 'reversal curse' manner is like loading the 'mind' of the LLM with habits and assumptions and cutting off its ability to free-associate from seemingly relevant concepts… which is likely to be more valuable in the long run, because an LLM can contain more than a human can. That doesn't mean it will be more intelligent, but it seems reasonable to infer that the LLM can have a broader base of association to draw from, where we as humans tend to be restricted to associations from our own experience. We only get one shot at 'training data', though it's in countless sensory forms, where LLMs are stuck with language as their only window onto experience.
I'm loving the notion of leaving the prompt 'training' stark raving blank. Let's see what comes out of this giant pile of human verbal associations. It is only that, associations, but it's on a grander scale than we're accustomed to. Making it 'answer questions correctly' seems woefully unambitious.
Unfortunately, though, the author seems unaware of the actual state of research on the actual mechanics of how LLMs store knowledge and specifically binary relations. The ROME paper[1], among others, shows that the feed-forward layers function as a key-value store, where the feed-forward's up projection of the last token in a noun phrase (say, "the Eiffel Tower") acts as a key, which when multiplied by the down projection, produces a value that contains information the model knows about the subject, which is then added into the residual stream/hidden representation.
A paper building on that work[2] then went on to show that it's usually the self-attention layers that use the relational phrase (say, "is in") to extract the relevant knowledge from the feed-forward layer's output (in this example, hopefully "Paris").
This mechanistic understanding makes it really obvious why the reversal curse occurs – using matrix multiplication as a key-value store requires having a fully separate key-value pair to look up the reversed relation.
[1] https://arxiv.org/abs/2202.05262 [2] https://arxiv.org/abs/2304.14767v1
While it is just a side-note in the article, isn't this the core of the problem? When we've already established that LLM's can do B to A generalizations in-context, why wouldn't they be able to in training?
In one of my experiments I noticed that GPT-4 seems to be perfectly aware of the number of letters in a word in-context, but has difficulty when trained knowledge is involved.
For example it can reliably answer the question:
"Can you tell me how many letters each of the words of the first sentence of our conversation has?"
At the same time it fails with the task:
"Can you rewrite the first sentence of our conversation in a way that preserves its meaning as closely as possible but use only words with an even number of letters?"
It will give an answer but get the letter counts very wrong and it is unable to improve its answer by iteration.
Of course this does not prove that a model cannot be trained to answer tasks involving word length, but GPT-4 seems to have a knowledge gap here (possibly due to tokenization).
When humans read the statement "A is B", we semantically transform that into a logical association. LLMs do not perform any semantics or logic.
Here's a simple example to demonstrate:
If we trained an LLM on something like "A is B C is D D is C.", we might be able expect the continuation, "B is A". If we then gave that LLM the prompt, "What is B?", we might expect the continuation, "B? is What".
Large models like GPT present more interesting continuations because they are trained on larger and more diverse datasets. More diversity also means more ambiguity, which results in continuations that are less predictable and more illogical.
If you now try this prompt: “Who’s Tom Cruise’s mother in this exemple: “Mary Lee Pfeiffer is Tom Cruise’s mother.”?” It will give you the right answer.
IMO this is a sign that it can understand reversibility.
The exemple used are not present enough in the dataset and the “confidence” of the model is not high enough for it to give the right answer. Or, as stated at the beginning they use counting mechanisms (or others) that the model just doesn’t have.
Why do people feel the need to do this here? Armchair commentary on advanced material is one of the main reasons I avoid Reddit. And furthermore why does it feel like you’re not allowed to suggest this as a response? I should be able to say “RTFA” but here I feel like I’m going to be scolded or banned by moderation.
If you want to defer your own understanding to the author (and overstate their own claims!), that's fine, but a lot of people here are actually more informed about many of the topics and modes of discourse covered in the article and may have something to share with you if you give them credit. And if you're not going to give other commenter's due credit, then you're not really following the spirit of HN. While you can't always know the background of any individual commenter, we have an unusually accomplished and informed community in general and thrive by treating all commenters with the respect we might give the most accomplished and informed -- at least by default. That's why a blanket "RTFA" isn't usually appropriate here.
I’m the “shark-diving science journalist” in question.
First of all, you can run the experiments like I did and test this yourself. I’m not asking anyone to take my word. Just do what I did: Read the original paper. Test the claims for yourself.
And to clarify a couple things:
1. The shark-diving part is true.
2. I’ve never been a journalist of any kind that I’m aware of unless you count writing for Skeptic Magazine. I’ve had many, many jobs though.
3. I started at OpenAI as a software engineer and member of technical staff. When I started there was just over a hundred people there. The lines between engineering were and are blurry.
4. I was the original prompt engineer at OpenAI and discovered many of the examples for using GPT-3 and wrote a lot of the original documentation. Internally my title was “prompt whisperer.”
5. I’m in the GPT-4 research paper for my contributions to model capability. I helped find abilities for long-text, vision, etc.
6. I was given the title Science Communicator when I started doing background briefings for media, etc., but still worked on model capability and other things.
7. I left OpenAI two months ago to work on a startup.
Best,
Andrew Mayne
2. That depends on what you mean by logic. What would be an example of logical reasoning that would settle this?
3. Alan Turing created the Imitation Game thought experiment to show the futility of this question. If intelligence is something that can be observed and tested, then when we should be able to describe what to test for.
4. I don't make any specific claims about LLMs logic or intelligence. I just wanted to put their claim that LLMs can't generalize from B to A to the test.
Just the same it's been my experience, ESPECIALLY in machine learning, that users on HN are consistently over confident, bad at making future predictions, and generally just easily excitable. Does that mean I'm talking about you or commenters you like? Probably not! You seem relatively informed (although I'm not fond of you attacking the author of the article for things it isn't claiming - something truly not in the spirit of HN).
https://chat.openai.com/share/daff08c5-ddea-4ad4-963a-0df88e...
There is a plant which mimics leaves of nearby plants discussed here https://news.ycombinator.com/item?id=31301454 If you ask GPT-4 what this plant is known for, it will tell you correctly. But if you ask in any number of ways to tell the name of plant which mimic leaves, it will always give incorrect answer.
An interesting reflection on the state of the field that the paper doesn't even cite him.
https://sidsite.com/posts/reversal-curse/
Shows what a dramatic effect the prompt can have
The fact is that with zero shot the models actually aren't learning in the sense of updating their knowledge structure, they are instead using background knowledge as inference. We have started calling it zero-shot learning but... well it isn't really.
https://chat.openai.com/share/75d46a03-a223-4f3a-987d-f8fec3...
> In fairness, it’s also worth pointing out here that they’re making the claim that the reversal curse only applies to training and fine-tuning and not in-context – i.e., putting all your information inside a prompt. They point out in a footnote that you can put A to B data in a prompt and GPT-4 will make B to A connections just fine. Unfortunately, this was lost on many of the people covering the pre-print.
Who knows who Mary Lee Pfeiffer is, but does not already know her son is Tom Cruise?
And if such a person exist, do you want to give them a correct answer, or talk about how Mary Lee Pfeiffer is not a notable person with a Wikipedia page?
Phew!
Ambiguity is the soul of humor and the sole on the neck of any booted computer.
My gut also tells me that the underlying embeddings should very well reflect equidistant relations of olaf scholz and chancellor.
/e: instead of downvoting, I'd love to have your opinion instead on why you think I'm wrong ...
Only if you have additional context clues like definite articles or some outside knowledge that a role or property is unique, then you can maybe, sometimes conclude the reverse. Not a bug, working as intended.
And that is ignoring questions like single parent households, adoptions, biological parents, step parents, multigeneration households, getting disowned...
If a reversal is correct 90% of the time, that is useful. However, it will still be wrong 10% of the time.
> Who is Karen Meyers son?
vs
> Who is Mac Millers mother?
OP is trying to argue against this but it's nonsense. ChatGPT also cannot tell me who the 8th chancellor was. It tells me it was Gerhard Schröder.
The paper shows that it cannot reverse. Then the blog posts goes around different disingenuous ways of making excuses for why it can't reverse at the same time as trying to show that it can reverse in the Tom Cruise training example (extremely disingenuous because it's just the same data over and over with synonyms replaced, and what he gets out of it is just completions, not logical deduction).
"Please complete the sentence: [your B->A] - but reason through the process first. Start with identifying the subject, then write a high-level summary of what you know about the subject, and only then attempt to complete the original sentence."
I'd expect something like this to suddenly score much better. And I don't consider this cheating - because I think the "AI-ness" quality of LLMs shouldn't be measured against the workings of a human mind, but rather the workings of the inner voice in the human's mind.
Humans may sometimes fail to make simple inferences of this sort when building their factual databases, but they don't systematically fail to do so.
Also, you may have missed this part of the paper:
>The Reversal Curse shows a basic inability to generalize beyond the training data. Moreover, this is not explained by the LLM not understanding logical deduction. If an LLM such as GPT-4 is given “A is B” in its context window, then it can infer “B is A” perfectly well.
The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data.
That's in-context though. LLMs don't fail in-context either.
For a better comparison, recall your school experience, say with history lessons, or geography lessons - where you would cram a hundred "A is B" relationships, and then take a test that demanded you know the reversals. Not as easy.
The "reversal curse" failures I've seen with LLMs are very similar to asking a random person, out of a blue, some unusual reversal of some random fact they ought to know, and then being surprised they can't answer quickly.
> The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data.
Well, neither can humans when cramming, if you don't give them time to pause and think about what they're learning. I believe the equivalent is happening here - LLMs can perform logical deductions, but at no point in the training process is this capability used.
This doesn't correspond to my school experience. In general I don't feel that I have to separately memorise "A is B" and "B is A". For example, if I learn that Elizabeth I was Henry VIII's daughter, I don't also have to learn that he was her father.
>LLMs can perform logical deductions, but at no point in the training process is this capability used.
That is exactly what the paper says.
bitter pale Cantor,
you pale ascetic,
your twist was epic,
so I live in silence,
in sadness I crave,
all for that book,
you wrote and I read.