Can LLMs learn from a single example?
fast.ai
fast.ai
I'm one of the authors of this post -- Johno & I found it really interesting looking into this curious issue of rapid memorization from LLMs. I've been working with neural nets for 30 years, and fine-tuning language models since 2017, and this behavior is most surprising to me! Other folks have seen it in LLMs too, although I haven't seen a analysis of this kind before (although we might have missed something).
Let me know if you have any questions or thoughts.
From an average -87.3% performance drop on the 12B model to -61.6% on the 84B model then just -3.9% on the 562B model. Felt like we were just shy of an insight breakthrough here.
Is avoiding CF potentially just a matter of sheer scale ?
So I'm not even sure we're showing any problem to solve here -- it might be more of a opportunity, in fact!
This scales well, too. There are facilities that provide services of co-hosting and cross-training up to ~two dozen NI models in a shared environment - in my experience, this provides similar training benefits to running multiple NIs on your own, at fraction of the cost.
(The facilities are exploiting some neat economies of scale. Talking to some employees, I learned that the transfer learning and co-activation are embarrassingly scalable: if you get two-three NIs to pick up a thing, all the rest immediately follow.)
The accounts aren't wired up by default to the AI and I am refactoring the templating system right now, but you can definitely start storing and searching things.
Is this 'overconfidence' the leading explanation as to why LLMs continue to show qualitative improvement even after their test loss levels off?
I assume this means losing all the energy and compute input for a model to know, perform, infer on inputs already indexed(?) (What is the proper term here?)
But is this the premise -you lose all prior investment of resource to a (I don't know the term for an AI archetype of knowledge) {btw, I love the embedded etymology of knowledge
"The ledger of things that we KNOW"}
It occurs because training changes the weights of the model. The earlier set of weights was good for the previous tasks. The new set of weights is only good for the new task. Usually special care must be taken to overcome catastrophic forgetting.
But imagine all LLMs in a macro view like a sponge entity
Note that the only reason that things are catastrophically forgotten, is that the original examples are not shown again. If the model learns in a single shot, there might simply be no time to show both the old and the new examples. I don't think it would have a significant effect or else we'd know about this effect a lot sooner (i.e. the training of these LLM's would get less effective from a certain point)
This seems like a ripe angle for evolvement of our understanding of AIs use in LLMs... can we throw AIs at AIs (is AI synonymous to LLM?) Can we throw LLMs at LLMs? and have them recursively learn from themselves.. or is it a Rat King. AI recognize AI in the GangPlane
In any case, looking at and understanding how a neural network encodes information is like gene editing. Perhaps you could isolate a gene in the human genome that achieves something interesting like giving a child blue eyes. But even if you would do that, there's a chance you break something else if you modify that gene and give the child health risk. Since all neurons in a deep neural network are interconnected, there is a butterfly effect in it that makes them inherently somewhat of a black box.
My intuition would be that you get more orthogonal directions to the gradient (of previous samples) if you have larger model.
Although I am not a researcher, it is obvious to me that not all LLMs are the same architecture, and I think that even ones with similar architecture can evolve to functionally operate quite differently on the same inputs.
Yet most articles seem to refer to LLMs as if they were just one architecture and model.
Is that not correct? (Of course I fully accept your expertise in this matter and this is just my curiosity, not trying to tell you you're wrong!)
> So, MOND reduces the discrepancy in clusters at these radii to only a factor of ∼2−3 (Sanders, 1999; Ettori, et al., 2019)
So I think you're right, and today I learned something! I also checked if Stacy McGaugh had weighed in on this particular subject, and it seemed like there is still an issue for clusters [2], although interestingly the issue isn't mentioned in his latest blog post that summarizes the strengths/weaknesses with MOND [3]. Anyway, thanks for humoring me for a bit.
[1] http://www.scholarpedia.org/article/The_MOND_paradigm_of_mod... [2] https://tritonstation.com/2021/02/05/the-fat-one-a-test-of-s... [3] https://tritonstation.com/2023/06/27/checking-in-on-troubles...
If this is so, then a multi-choice question which conflates one particular MOND theory for MOND itself, and which depends on the specifics of that particular theory for selecting the 'correct' answer, is problematic: for one thing, it may make selecting the 'correct' answer more difficult for a student who has specific knowledge about the topic. This is just one of several problems with multi-choice questions, though, fortunately, it does not seem to have any bearing on the very interesting phenomenon you have discovered.
I've noticed somewhat similar behavior while training graph neural networks to model physical systems, except that it takes way longer than a single epoch to get there. Or course, there's no pretending involved with my GNNs, but the models do have very constrained representations, so once they start to figure out how to represent the physics at hand, the loss plummets dramatically.
Like online learning in a way. But you do it during inference time.
There’s no way the entire model actually needs to be touched for something like “sky color is:” and “blue”.
Drop a set of neurons and there’s no change? Probably doesn’t contain the “sky color” concept.
Drop a set of neurons and the model freaks out, definitely conceptual neurons.
Rinse and repeat to find the distilled pattern across all the neurons.
You could train an LLM against the neuron graph to do this for you.
From a physics perspective it’s entropy: There are just more local minima that have neurons code multiple things.
I suspect that dropout and similar tricks also play a role in this: Removing connections during training means that pathways need to be redundant somewhat.
What is happening is called “over fitting”.
Think of data as dots. A model that generalizes well will create as simple of a function as possible that fits the training data points pretty well.
But keep training and parameters will often get very large, creating huge up and down swings in the function curve, far outside the actual data values, in order to pass through the training data points exactly.
So it’s technically a better fit to the training data, but it is now a crazy function, often producing extreme outputs on new data. Practically a worst case lack of generalization.
Thus, “over fitting”.
And “over fitting” isn’t the same as “memorization”. Large models can memorize small datasets without over fitting. They have so many parameters, it takes few changes to fit the training data. At which time, learning stops at an otherwise random function, and generalization is never achieved.
That case is called “underdetermined”.
There are models that produce both outputs and confidences (essentially predict their own error standard deviation per output, based on the input).
So “over confident” can mean a model that predicted high confidence (low error deviation) inaccurately.
The issue here is one of calibration: https://en.m.wikipedia.org/wiki/Calibration_(statistics). That is, the output probabilities of the neural network do not reflect the true (observed) probabilities. If it is systematically underestimating the probabilities, it is termed "underconfident", and if overestimating the probabilities, "overconfident".
Note that in these cases, it may still be improving as a classifier on unseen data, while still showing higher validation loss as calibration degrades.
This is many decades old terminology for a well established effect that occurs for all curve fitting, function approximation, parameter optimizing, and model training algorithms.
You can Google it with no other context: “over fitting”. [0]
“Confidence” isn’t its name and its meaning has nothing to do with the effect.
Nothing wrong with making up terminology for new effects, but this one is an oldie.
[0] https://www.google.com/search?q=over+fitting&ie=UTF-8&oe=UTF...
In both cases, training to a perfect fit of training data makes no sense.
Learning the particular noise in training data guarantees worse results on real data, where the noise will, by definition, be different.
The same is true for exactly reproducing data that depended on some unaccounted for information. Once applied, the unaccounted information in the training data will have no predictive quality.
So perfectly fitting data is usually a terrible idea, even when it can be done.
Training data captures a problem to be solved, as best it can. But it isn’t the same as the actual problem.
——
Best practice:
1. Use training data to optimize a model
2. Use separate validation data to approximate current generalization quality, and stop at (or revert to) the point where validation performance was best.
3. Use separate test data, not used for model design in any way, for a completely independent approximation of generalization performance.
Wide disparities between validation & test performance suggest problems. Ideally they should track each other fairly closely.
If not, more data is probably needed to characterize the problem more reliably.
NEVER just retrain, or tweak training parameters, until training & the validation stop produce good test performance. That means the test data was actually used in the design and is no longer an independent measure of performance!
——
Assuming you are having trouble getting similar validation & test performances:
A good way to ensure test performance is real, is to retrain the model in exactly the same way, several times, with different random divisions of training, validation & test data. If test results are good regardless of data divisions, then they are reliable.
Then you can eliminate the test performance dependency, by randomly selecting any of the models. Don’t choose the one with best test results!
That takes discipline!
In that case, the mean of all test performances is your best estimate of generalization performance, regardless of model.
(Throwing out models with the worst and best test performances is ok, it avoids outlier training failures or successes equally. Note I said “avoids”, not “eliminates”, since the worst & best test performances are estimates of generalization, not a measure of actual generalization.)
What users want is to have their own credences properly calibrated by engaging with some system. From a physics textbook, they want a systematic presentation of ideas which allows them to build intuitions etc.
It's important to formulate the actual goal of the system, rather than just the engineer's goal (consider eg., "width of pipes" vs., "clean running water").
In the case of statistical AI systems, the goal is often best formulated in terms of the confidences of the system not its output. Since its output accuracy is kinda nonlinear and discontinuous in those confidences.
So from a statical AI Q&A system we dont want The Answer, we want the system to have expert-like confidences over possible answers.
Of course, as soon as you start formulating these metrics, all the SoA 99%+ accuracy hype evaporates. Since most of these systems have terrible confidence distributions.
Consider, eg., ChatGPT whose answers are often plausibly accurate (they count as an answer) but just repeat some silicon valley hype in a way an expert wouldnt. ChatGPT rarely has the careful scepticism of an expert, rarely presents ideas in an even handed way, rarely mentions the opposite.
It makes generating reference materials on areas with expert disagreement quite dangerous. ChatGPT presents the non-expert credence distribution. (And indeed, always does, since it just models (Q,A) frequencies which are not truth-apt)
I'm saying as a matter of fact ChatGPT should have different confidences in propositions. My issue isnt the tone of voice, my issue is the content of what it's saying is wrong wrt what we care about, ie., expert credences (/confidences) in the claims it's generating.
It can "express confidently" scepticism; it does not. That's the issue.
In my lang above i was mostly using credence to talk about the strength of the mental state of belief; and confidence to talk about the model of that used in statistical AI.
I do think it’s a form of overfitting, but a weird one. Overconfidence seems like a good, more specific term to me.
You have a generative model with billions of parameters that already assigns some probability mass to your (fine-tuning) samples. Now you compute a gradient that increases that probability mass, and take a step in the gradient’s direction. Essentially the OP is surprised that this significantly increases the probability mass of the samples under the model.
I’m not very surprised. The generative model is enormously over-parameterized and already assigns some probability mass to the (fine-tuning) samples. It would be surprising to me if there wasn’t a direction in this billion-dimensional parameter space that rapidly increases the probability of the relatively few samples.
ie. if they are only being trained from one epoch, there is clear overfitting concerns just by doing even a second pass in the data.
It does seem somewhat contrary to the findings of this paper [0] that found that old data was as good as new for at least 4 epochs.
Slight nit: Many public LLMs are trained for at least slightly over one epoch, and usually several epochs on particular subsets of the data (like wikipedia).
> Also Meta team with llama show that simply training more, more tokens, continues to reduce loss.
Can you source the specific claim you are talking about? More tokens to me generally will mean new tokens unless you are specifying.
from the paper "We train for one epoch over the training data. In earlier experiments, we found that training longer can lead to over-fitting"
You can clearly see in table on second page that higher quality data is trained on more than 1 epoch. Most Open LLM's do this.
If you are talking about GPT-4, unless you are insider (doubt it) you'd have no way of proving it either way because that info is not public.
> Llama2 doesn’t and outclasses llama.
My point still stands, llama2 is just one llm, and we still don't know the distribution of their training set.
Nothing it tried worked, it got close, but it didn't work.
Finally I found some C# code that fixed the problem, and I pasted that code into ChatGPT, asked it to read it, and then fix the problem in PowerShell.
It said it understood the solution, updated the script, and it worked perfectly.
For some reason that behavior was pretty eye opening. Providing material in the question that it wasn't trained on made it solve it.
It's understandable how it did it from language training, it just felt very cool that LLM's can do that.
These things are really easy to anthropomize, partly because they are good at "talking" and "articulating". So good that we tend to just accept that magical, enormous feat of statistical engineering as a trivial building block. But it's a brick made of gold.
Translating (from natural language to code, from text to audio, from image to image, one natural language to another), editing, summarizing, expanding/extrapolating is what these models do.
The inherent "knowledge" is just context.
(1) Vector embedding is in my view a little different - it's a form of semantic cataloging (akin to Dewy decimal) - and certainly enables search.
But "data retrieval" (who was us president in 1984) directly from the models isn't really all that interesting IMNHO.
I wonder what would happen if you trained an LLM on a little input but then had it generate a lot of synthetic input added to the training data. I think of it as "dreaming". This seems like it would just add noise, but LLMs are able to improve their output by augmenting their own context (by "thinking out loud"), maybe they can do the same with their own training data?
That said, it's common to use large models to generate synthetic training data for training other smaller models. In this way, we're able to transfer knowledge from one model to another.
Isn't learning from a single example desirable, while memorizing undesirable in the context of training? The former is the goal we're aiming for in order to match how animals learn, while the latter a failure mode that happens often. The article shows a case of unexplained memorizing, not of learning, right?
I haven't trained convnets at this scale so I'm not sure if similar behavior has been seen there, but you'd think someone would have mentioned it at some point. So perhaps these strange loss curves are a feature of Transformer based models in particular?
this is basically the case with transformer networks, which is apparent when learning from scratch. The model seems to be going basically nowhere and totally useless until suddenly, at some random point after a bunch of learning cycles the weights find some minimum on the error surface and bam, suddenly the model can do things properly. And it's because the transformer has learned an abstraction that works for all of the input data in an attentional sense (think how you scan a sentence when reading). Not the best explanation but its from memory from a post I saw on HN a while back
If memorization of events in llm is accelerated because of- these deep semantic frameworks, then does this provide a path towards long context windows?
I like the idea. You would need your own mutable copy of the model, which is usually huge. And you need to backprop so there is a bit more computation. It might be doable for a local model that is smaller than GPT3.5/4.
You also need to decide what is worth memorizing long term vs short term.
It could just be the diff against the main model or similar.
If your domain is heavily algebraic, you might even be able to generate correct examples arbitrarily, which is a situation I recommend anyone to be in.
This metric can be learned so it’s okay if it’s really hard to specify.
[1] https://memit.baulab.info/ [2] https://rome.baulab.info/
Overconfidence bugs moe because if you want to turn predictions into decisions and actions you have to be calibrated. I’ve found that some of these models that look like they are over fitting on loss are actually still improving on AUc (matters to me more than accuracy) and I can put a calibrator after the model to get the results I want.
(Still, for my current problem which has noisy labels, I find embedding + classical ML performs as well and takes a fraction of the time as fine tuning and clearly shows benefit trained on more examples than FT does. If I was going to do more model engineering on this problem I would probably resort to “stacking”)
The report mentions there is no reshuffling: > We’re not re-shuffling the dataset at the start of the epoch, so those first batches of the second epoch are when the learning rate was still warming up.
It is not: https://en.wikipedia.org/wiki/Belief_perseverance
> I'd definitely update my knowledge after a single failed test question
Maybe you would, maybe you wouldn’t. There are several psychological experiments which show people don’t act the way they say they “definitely” would when confronted with the situation. Quite a few examples in the movie “Experimenter”: https://en.wikipedia.org/wiki/Experimenter_(film)
> if it was something I'd care about, and I discovered my previous model of reality was wrong.
Those two ifs are doing a ton of heavy lifting. LLMs neither “care” nor “discover”. It’s not like you’re giving it a new contradicting piece of information and it’s going “interesting, let me research on that and update my model of reality if after careful consideration I find your assertion to be true”. It’s closer to having someone who’ll just accept everything you say and repeat it.
It might be more effective to try to play 'tokendle' before trying to play 'wordle'.
Or would an LLM get confused if we were to alter the way the tokenization of the input text is done, since it probably never encountered other token-"spellings" of the same word?
https://www.geeksforgeeks.org/lzw-lempel-ziv-welch-compressi...
For 'code table' substitute 'token table'.
It's basically unrelated to what happens during training, which is using gradients.
Sure. Neural nets in general can: after they've been trained on billions of examples first.
It really helps if they've previously seen the same or similar "single example". Which, let's be fair, the larger the training data, the higher the chances they have.
>> This seemed, at first, quite impossible. It would imply that the model was learning to recognise inputs from just one or two examples
To be more precise: the article is talking about fine-tuning a pre-trained LLM, so that's a-few-billion-plus-one-or-two examples.
Btw, what model was that? The article doesn't say.
Why do you say so? We casually call it "connecting the dots". It's like during the Oppenheimer movie when after the first demonstration of Uranium splitting people thought "oh, we can do a bomb with that".
Yes, ideally the former will associate well enough with the latter that, once you find some reason to think about mRNA, it will automatically drag up the thing you learned earlier and then you'll update. But it doesn't happen by itself, and sometimes it doesn't happen at all. Most people contain significant inconsistencies -- I would dare to suggest most likely everyone.
Essentially it understands programming didnt know what was possible in angular16 A single example made it learn from it. Though when i asked for an example i got the exact same sample as i had given it to learn from.
Perhaps end this language cut of for technical data. Its okay not wanting to get into politics (neither do i). But give it something to read (yup let it read and remember it) a simple prompt read this page by page will do, and give it some recent books, or popular coding websites, let it read python.org angular.io perhaps some modern manuals and books.
It also seemed keen to learn new information, it quickly adopted it. But only in that session.