It seems like understanding when "is" is reversible is pretty core to the capabilities of the model, but that's different than having a lot of facts memorized. For instance, a model should be able to answer "who is the star of Mission Impossible" with "Tom Cruise" based on context or internalized-training of "Tom Cruise is the star of Mission Impossible" while not answering "Tom Cruise" to a much more general question like "who is covered in water" even if it had internalized a bit from a review like "there's a scene in Mission Impossible where Tom Cruise is covered in water..." unless it had context pointing it at that specific case.
Best to simply scale log-likelihood-based training, next-token-based training trivially contains a requirement for learning all of the subproblems that predict said, next token, and hardcoding something to get warm human fuzzies would be creating a biased estimator (and move us back towards the 90s a bit).
Models already very constantly do context-dependent token utilization, it's an autoregressive feature based on the entire stream of incoming tokens. Humans have a bias to focus on the 'last token used', this is not what language models look at.
But human language is created for and by humans. Is not then operating on language in a categorically different manner an incorrect usage/understanding of language?
For more information, please see https://people.math.harvard.edu/~ctm/home/text/others/shanno...
The version with the Weaver introduction is quite good as well, there are other versions of similar papers covering the topic from different angles, I find it to be well-worth the read.
But are they? Unless I'm misunderstanding what you mean here, I'd say the opposite is clearly the case - learning "A is B" doesn't automatically mean learning the reversal; you either have to learn it from external source, or infer and memorize it - both involve an extra cognitive effort. This maps to LLMs as training, or "in-context learning" and then adding the inference to the training set.
In the non-simplified case I'd actually say that we have some natural experiments, but that might spoil this as well because you can in fact say that this is training for such invariance (which I'm unsure is problematic to my claim but pushback is definitely welcome). If we're learning history we'd often learn information such as "On December 7th 1941 Japan attacked Pearl Harbor and this was a major event that got America directly involved in WW2." But then on the test we'd see questions such as "What important event happened December 7th 1941?", "What was the date of the Pearl Harbor attack?", "What major event led to the US becoming a direct player in the war and when did this event occur?" and so on. Granted, because we do such testing I'd expect GPT to be able to reasonably answer many of these same questions but I'm very much under the impression that humans are far more robust/generalized and would quickly learn these reversals. And to stress, conditioned on knowing the forward direction to a strong degree. If you're fuzzy about the forward direction I'm not surprised if the reversal is even more fuzzy. Definitely biased. But more importantly, I do not think we have the same strong algorithmic bias in our encoding methods that transformers have. That's my main claim, just to clarify.
But I'm conjecturing a lot here. I think these are reasonable and we have some strong evidence for this. Maybe this is more situational than I'm presuming and I'm definitely working off of biased sampling. So discussion helps refine the ideas.
I think anyone who's used Anki to learn anything will tell you reversal is by no means free.
고양이 on the front deck, Cat on the back.
You can see this direction and remember cat is the correct answer every single time but completely blank in the other direction. Cat on the front. What Korean word is at the back ?
Conditions I mean: 고양이 => <concept of cat>; cat => <concept of cat>; either [고양이 => cat || cat => 고양이]. Under these conditions 고양이 == cat && cat == 고양이 should be known.
Your conditions: 고양이 ~> <concept of cat>; cat => <concept of cat>; [ 고양이 -> cat && cat ~> ...고양이?] (or something, trying to express you're preferentially learned one direction, usually mapping back to mother tongue while learning)
By no means are we strictly invariant. But I'd say that a good measure of your skills are in fact to be invariant. But also like you are pointing out, we also kinda train ourselves that way because it creates stronger pathways. Training methods are important too! And augmentation is critical to both human learning and machine learning. They're specifically critical to generalization abilities. You'll also see this in humans. I notice test focused educations tend (there's no "always") to create less generalization of concepts as you're over learning a thing presented in a specific way rather than focusing on the concept as a whole. Form of metric hacking if you will.
Side note: if you have tips for an extreme beginner learning Korean I'd really appreciate them. Trying to reduce the burden on my gf having to translate everything for me when we're around other Koreans.
If you have learned Hangul, then you can start with Lingodeer(app), TTMIK(talk to me in Korean)’s Grammar books, or Billie's beginner playlist - https://www.youtube.com/watch?v=sx0yyQqkpqo&list=PLbFrQnW0BN....
These are all extremely beginner friendly. You can just look at the options and choose whichever works best for you. I personally tried lingodeer and ttmik both. I did a bit of lingodeer and then ttmik levels 1 to 3. I didn't try Billie's playlist, but he's a great teacher, so I recommend it if you think that would work better for you.
So in the beginning stages, it doesn't really matter so much, but eventually you're going to have to start thinking about how you’ll be learning the big 6, vocabulary and grammar (what I’ve spoken on so far covers grammar and TTMIK takes you pretty far in grammar if you stick with it) then speaking, listening, reading, writing, because believe me, there isn’t as much transfer learning as you would expect or hope. I mean, you don't have to stress over it or anything, just keep in mind that you have to eventually think about how you're going to address those things, and the sooner, reasonably, the better.
So this is kind of where my training diverges from most traditional approaches. I know I used Anki from my example and a lot of learners use that with pre made decks from the community for vocabulary, but honestly, I just couldn't use it then.
Part of it was because at the time, learning Korean words was like drawing blood for stone. For me, it was so hard. Even five words a day could be a challenge. My brain was like, “what the hell are you doing?”. It just wasn't going in.
Come to think of it, you know grokking in NN’s? Learning a language, you quickly realize the same phenomenon can occur in our biological counterparts. It’s like one day I was struggling to learn these new words, I don’t do anything drastic or different and suddenly the next day, learning new words felt so much easier, it’s crazy. Progress isn’t necessarily linear that’s for sure.
Part of it was also because I just didn't like the Anki experience, memorizing hundreds, thousands of words in isolation. It just wasn't really doing much for me. So, I just kind of dropped it. If traditional anki works for you then go for it but gin, you’ll need to figure out the remining 4. No matter how much isolated grammar and vocabulary you learn, you’ll need to practice doing each of those things for it to sink in.
The typical process is to learn a few thousand words and up to intermediate grammar before sinking into the 4. Makes things a lot less frustrating.
What I found was this amazing audiobook series. It's a graded reader series. 100 books with audio, all free(start from the oldest). https://audioclip.naver.com/channels/57.
Here the books are further divided into stories. At first, each story is really a few pages but it does get longer and longer as the series progresses. graded readers are specifically simplified literature created to help language learners progress.
It's also so that you can read in your target language much earlier than you otherwise would have because, well, you might see a lot of people say, “oh, just read, just watch or read children's shows or children's books.” That's a trap because like, children are native speakers man. Don't make the mistake of thinking that they aren't.
The complexity of the stuff they read or watch or whatever, is well beyond the beginner learner. So if you go into children's media and learning some few hundred words thinking, “oh, this should be easy, then, yeah, it may become very demoralizing because it won’t be in fact easy”.
And just to clarify, I mean children’s stuff not baby stuff. Yeah, baby stuff is simple, but you're going to get bored of that very, very fast.
So, yeah, it's a graded reader series and the idea is just to get started. Simplified stories you might find more interesting than baby stuff.
So, I really like reading. And I don't know, I just felt like it would be better because, you see, what Anki does is SRS, spaced repetition system, right? But when you think about reading, it's really a natural form of that. So, I just felt like, okay, this is what I'm going to try.
So, I was even thinking, oh, I could use this to knock almost all the birds with one stone, you know, it’s vocab, grammar, reading, and listening practice at the same time.
But the book series, it starts with a pretty high level. I don't know if you're familiar with the CEFR levels. It starts at about a low B1(bottom end of low intermediate) or so and ends at a high C1(top end of low advanced).
So, yeah, coming from where I was, where I knew at best, maybe 100 or 200 words and jumping in, it was difficult. And honestly, I would not have been able to do it without Mirinae. So, Mirinae is a grammar parser - https://mirinae.io/.
And what it does is, so basically, Korean and English are so far apart that you could know the meaning of every word in a sentence, and you would be absolutely lost as to the meaning of the sentence itself. So, what Mirinae did, or basically took what would have been impossible for me and made it achievable. I learnt a lot of grammar this way. You place the sentence in, and it would basically parse the grammar on the sentence.
You could see the explanations of different grammar parts, different nouns, objects, stuff like that, how this affected this and all that.
It was very, very, incredibly helpful. To to this day, I think that at least for distant languages, a good grammar parser is by far the most useful tool, better than a dictionary, better than just a translator.
So, the first story, I was pasting every sentence. So, it took me about a week to get through it, a few hours every day. The amazing thing was that you could the progress, the reduction in time it takes you to complete things, because the second story took only a few days in comparison. By the third or fourth story, I think it took only a few hours. I mean, I didn't finish it in a day, but I could have if I wasn’t a bit lazy.
So, yeah, that's the benefit of graded readers. Whether you start this as early as I did or not, it’s genuinely an amazing resource. Earlier this year, I came across this Anki deck series with natural audio -https://ankiweb.net/shared/by-author/374470252
What the guy managed to do was, he took a machine learning Korean speech dataset, and he kind of configured a lot of the sentences, which were, you know, a single woman speaking sentences, about 12,000 or so of those sentences. So, he reduced it to about 7,000, and he rearranged, as much as possible to be i+1 comprehensible. i+1 is just basically the idea that you come across a sentence, and there’s only one word you don't know.
So, it's much easier to grasp the meaning of that sentence quickly and easily. Then the next sentence, only one word, then the next sentence, only one word, then, and so on and so on. Now, this deck isn't perfectly i+1. But going through it, it's very good. So, you know, you could also explore that too after the beginner grammar resources, I suggested.
I’ve also noticed these guys, https://umiapp.co/.ly. It’s an app tailored for learning words in context, comprehensible input style. I really liked the looks of what I saw from the other languages. Korean isn’t there yet but is coming soon. By the time you’re ready for it, it’s hopefully there.
Anyway, best of luck.
Similar criticisms are "well there are many descriptions of game x it would have read so why doesn't it play x well". But gradient descent is a dumb optimizer. Training itself is not actually like someone reading a text anymore than evolution is like someone thinking about the best way to augment an organism.
a "smart" optimizer would look at a reversal applicable sentence and know exactly what bunch of weights to change to store it in such a way as to be recalled reversibly in the future.
Inference may be smart (GPT-4 can reverse in context just fine, potentially play games from only a description) but the training is not.
> You could potentially describe how a game is played to GPT-4 without examples and get it playing it correctly but passing that same description into the training process of the model just gets you a model that can describe your game correctly.
It seems like you're referring to an associative Hebbian/Hopfield-like lookup, which the current 'dumb' optimizers already do. Better yet, the learning rates are normalized by the diagonal of the empirical Fisher so that the learning w.r.t. to some estimated expected information is more constant, for said associative lookup operation.
Additionally, the training loop (which you call 'dumb') is just...a teacher-forced version of inference, which you call 'smart'?
It's better to simply minimize the log-likelihood in a scalable way. Hand-engineered solutions rarely survive compared to strong-scaling ones.
>Additionally, the training loop (which you call 'dumb') is just...a teacher-forced version of inference, which you call 'smart'?
What powers In context learning is not very well understood but it doesn't appear to be or really work exactly like just a non teacher-forced version of training. There are qualitative differences. The same models have no problem with this 'curse' when the information is provided in context for example.
>It's better to simply minimize the log-likelihood in a scalable way. Hand-engineered solutions rarely survive compared to strong-scaling ones.
I never said anything about it being bad.
I mean, yes. One distills information from a training set into a compressed representation, and the other generates a compressed representation that yields (more or less) fixed state space attractors. It's just inducing a bias over the state space of the network, nothing incredibly special, though I'd consider the initial stage of training to be the most important, as it is responsible for all of the ingest of all of the embedded information the network will be using during autoregressive inference (especially w.r.t. the context of doing so during longer generations).
So the notion of 'in context learning' is one I find to be a bit of an illusion, of course, as no actual learning is being done, just the induction of biases, which appears to give rise to a transiently-'better trained' network.
You could see this as a bit straightforward perhaps, but I feel it needs to be said.
(Perhaps though it has to be done during training so that "Who was the ninth Chancellor of Germany" becomes knowable in the first place.)
No, in English this would be “apples are red”, not “the apple is red”.
"An apple is red" does not imply that "Red is an apple". Use of the indefinite article makes the non-equivalence clearer.
Not to Bill Clinton this, but I think you are considering this for the wrong interpretation of "is". In this case, "RMS is the founder of the FSF" is the sort of statement implied here. Nobody else was or ever will be the founder of the FSF. So yes, for what the author is describing, "B is A" should always hold. B, in this case, isn't a category, it's an exact identity. The author doesn't explicitly state this, but it's clear this is what they're talking about.
This problem exists for both LLM's and humans, I think it may be fundamental to reality itself, as it currently is at least.
I wonder if they changed the training data to consistently use "equals" where possible (even if it sounds weird to the reader) would make a difference.
The fact that the current ones don't is surprising.
I'm not going to hype the technology more than it needs, but it seems to do a pretty damn good job of just this. And to the author's point, it does get this right.
And in fact, I'd suspect many of the cases where the LLM gets this wrong are caused by lots of examples in the training data which were written at different times, where someone else was chancellor. Or where the LLM gets numbering wrong (which is a well established problem).
And this is just one thing that makes sorting out what's going on here difficult - essentially, we have biologically LLM's [thinking they are] decoding silicon LLM's, of course the results are going to be weird and counterintuitive.
[1] Examine the specific language you are using, but also the colloquial ~interpreted/virtualized assertion that one is left with, if they are reading your comment "in good faith".
They implicitly recognize (and model) patterns in text, then continue those patterns.
While the patterns that an LLM has modeled may align to the high abstraction we call language, LLMs actually work at a much lower level of abstraction: plain text.
It's the content of that text that is significant. We encode language into text, and LLMs infer something "close enough" to language by blindly modeling it.
Unlike humans, LLM's can consistently realize (or at least claim to realize) when they've made a mistake (like assuming one's beliefs are necessarily correct, [because some persuasive story]).
The entropy gotta get slung around somehow.
(1) Mortal was Socrates.
(2) All humans are mortal.
(3) Therefor all humans are Socrates.
Or in short: LLMs will handle "The apple is red. Red is the " just fine, because the first part is heavily biasing the answer. They won't often complete "Red is the " alone with "apple", because why should they?
>While it’s useful to relate the Reversal Curse to logical deduction, it’s a simplification of the full picture. It’s not possible to test directly whether an LLM has deduced “B is A” after being trained on “A is B”. LLMs are trained to predict what humans would write and not what is true (Lin et al., 2022). So even if an LLM had inferred “B is A”, it might not “tell us” when prompted. Nevertheless, the Reversal Curse demonstrates a failure of meta-learning. Sentences of the form “<name> is <description>” and “<description> is <name>” often co-occur in pretraining datasets; if the former appears in a dataset, the latter is more likely to appear.4 This is because humans often vary the order of elements in a sentence or paragraph.5 Thus, a good meta-learner would increase the probability of an instance of “<description> is <name>” after being trained on “<name> is <description>” . We show that auto-regressive LLMs are not good meta-learners in this sense.
Correction. The Model can. Training cannot.
It's not that humans aren't suffering from this too. It's possible to spot it when you're learning, or even easier, when you're teaching someone a new thing. A person who learned that "A is B" does not automatically learn that "B is A"; they need to first process it, perhaps run the reversion explicitly in their head. I'd say that checking if the student can infer "B is A" having learned "A is B" is a good way to tell whether they're starting to comprehend the material, vs. just memorizing it.
This isn't surprising if you think about how the machine actually works rather than treating it like a sacred magic box.
The default assumption for why there is any "success" for in-context learning should be that it's just picking a nearby token that "fits" not a process of logical deduction.
edit: The obvious band-aid-fix is to feed reversed sentences into the training data without telling anyone, after which LLM boosters will proclaim that LLMs "learned" to reverse logical implications.
I don't know in what way this is some trivial "recall" problem as you suggest. Does the reversed implication in question need to be in the training data explicitly or not? If it does I don't understand how you can claim a logical inference capability.
No that's not what the paper says. It says training won't immediately store this information in a way as to make it reversible.
Let's get one thing straight. This is a common problem for human learning as well. Anyone who has used Anki for language learning will tell you that if you just train on target language word on the front and native language word on the back, you will fail the reverse unless you specifically train for that.
This is specifically a problem of recall. Not a problem of making the logical inference. If you ask the human language learner immediately he has learnt the new word for the reverse direction, he will obviously tell you the correct word.
But later he may not recall the reverse even if he remembers the original. Again this is a problem of recall rather than the ability to make logical inferences.
In the same vein, if you give the LLM the original direction in context and immediately ask the reverse, it will correctly tell you. It can make the logical inference.
And although it's a complete waste of time because the answer is obvious when you do do these types of adversarial experiments, like adversarial blockworld, you find that LLMs are not internalizing abstract principles inherent to reasoning and are instead doing something like string-pattern-patching.
If GPT4 rarely makes those errors it's clearly not an insurmountable problem at this level of training, but it does imply a lot more difficulty fine-tuning models to reliably retrieve correct information from specific text.
I don't know that I even want these things to generalize. I get it in academic terms, but from a CTO perspective how much this matters to the business? We already know our policies and procedures and aren't exactly excited about the idea of something getting curious with flipping things around.
I don't want it to generalize and start making weird assumptions like "some men are tall". I want it to say "I don't know", or provide a statistical signal indicating the same.
LLM's have a specific architecture that makes them good at approximate retrieval. They have a specific learning process that makes them able to do this over a vast training set.
Doing things like reasoning about facts is possible using either symbolic or specific sub symbolic structures, but combining these with an LLM is a bit tricky. It's something that could be done for a demo in a 6hr/day/week project (depends on the demo) but to do it for real in an application... or as a robust bit of science... well... 6 people for 6 mths/years/?
The opposite statement, humans don't know when "is" means equal is also a true statement, perhaps even more true.
I doubt MS would put their service in lot of their products.
I consider them the target audience for upgrades, I think the ROI could be massive.
Can you explain what you are skeptical of? There is ample evidence that LLMs are what they say they are: a series of transformer (usually) functions with learned weights that can successfully be used to generate human like text in many domains. Do you disbelieve this? Or are you skeptical of something else.