Do large language models need all those layers?
amazon.science
amazon.science
I thought that it was common knowledge that LLMs are undertrained, none of the publicly available loss graphs show any sign of convergence!
This paper shows how the loss decreases when you increase the model size, compute, or training dataset size.
From the article:
> Convergence is inefficient: When working within a fixed compute budget C but without any other restrictions on the model size N or available data D, we attain optimal performance by training very large models and stopping significantly short of convergence.
It clearly states that when you are limited by your training time compute, you should under-train your model.
[1] https://arxiv.org/abs/2203.15556 Training Compute-Optimal Large Language Models
[2] https://arxiv.org/abs/2302.13971 LLaMA: Open and Efficient Foundation Language Models
- Pruning
- Distillation
- Sparse transformers
- Mixture or experts
- Quantizations
My understanding is that none of these are free and they all come with various trade offs.
For example MOE lets Mistral beat similarly sized models, and the inference performance stays close (only an incremental increase). But the training time is way more than a typical 7B model.
But which of these approaches gives the most bang for the buck?
Also consider it’s not either/or many of these techniques can be combined.
And maybe worst of all, some of the testing that can be done to find out doesn’t give the same answer with smaller/toy models.
Also it uses 12.9B parameters per token, not quite comparable to 7B models.
- Lookahead decoding
If we can trim a model in different ways to get different specializations, that could be really effective.
there's a finite number of relevant language tokens. the trick of current LLM is finding what's basically the center of a vast series of probability.
This also has become increasingly obvious from recent developments in the field. Today, we regulalry see new models that come at a fraction of the size of GPT-3 and yet easily outperform it, especially on certain downstream tasks when fine-tuned correctly. These small models also retain some generality, but not as much as the really high end really large models like GPT-4. I'd say a sub 10B parameter model equal or better than GPT-3 overall is achievable, but not for GPT-4. At least not with current technology. However, that would still imply that it's possible to reduce parameter counts in common approaches by 95%. I'm pretty sure in a few years people will look back and smirk at the crude methods we used to train LLMs today.
GPT4: 4 * 1.76T = 7TB
That feels too much, but I have no idea tbh.
IDK...
Nevertheless, it would be cool if it could be reduced 95% in the future, haha.
Most claims of models being even equal to GPT-3.5 are also significantly overblown. I haven’t seen one yet below 70 billion parameters which even comes close.
Nothing is free in this world.
really what we're looking for is a way to bootstrap appropriate filters so a unit of model matches a unit of language scope.
An 2022 paper "Scaling Language Models: Methods, Analysis & Insights from Training Gopher" (http://arxiv.org/abs/2112.11446) has captured it well on page 103, Appendix G:
> The general finding is that whilst compressing models for a particular application has seen success, it is difficult to compress them for the objective of language modelling over a diverse corpus.
The appendix G explores various techniques like pruning and distillation but found that neither method was an efficient way to obtain better loss at lower number of parameters.
So why does pruning work for OPT-66B in particular? I'm not sure but there is evidence that OPT-66B is an outlier: one evidence is in the GPTQ paper ("GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers", https://arxiv.org/abs/2210.17323) that mentions in its footnote on its 7th page:
> [2] Upon closer inspection of the OPT-66B model, it appears that this is correlated with the fact that this trained model has a significant fraction of dead units in the early layers, which may make it harder to compress.
Since this article specifically target OPT-66B and not anything else, I would remain skeptical until they generalize their findings to other language models, such as Llama2 or Mistral.
Chomsky famously said "Colorless green ideas sleep furiously". An LLM can detect that such a sentence, such an interpretation, or translation, would make no sense --even if syntactically proper-- because its training materials never include such sentences (except perhaps this particular one). Then knowing it cannot be that "Colorless green ideas sleep furiously", it can prune out many otherwise possible interpretations. This is just my hunch, something like this may be going on.
https://en.wikipedia.org/wiki/Colorless_green_ideas_sleep_fu....
I wish arbitrary topology networks were scalable (I love NEAT) but bipartide graphs crunch good in GPUs.
> ...if we meet the shape 512x512 (512 is a power of 2, therefore a very "round" number in computer science), then maybe some kernel will be very fast; but then on a 512x511matrix, the same kernel may need to add some padding first to transform it into a round 512x512 matrix with zeros at the end of each row. Adding those zeros means shifting all rows, which is a very costly operation.
https://arxiv.org/abs/2212.09095
since it is at its one year anniversary, we can check citations - it has 16 so far
https://scholar.google.com/scholar?start=0&hl=ro&as_sdt=2005...
But anyway it made me wonder if there's a way to measure "what x% of a model is actually used" similar to the myths about human brains.
That being said, only around 2% of the human genome is coding, with perhaps another one or two percent having known no coding function. And estimates of functional DNA is somewhere around 8% of the genome. So, most functional DNA is still unknown.
Am I interpreting you correctly if I say: "Finding the slope (training) may require those extra layers but finding a particular y value given an known x coordinate (inference) may not require those extra layers".
What I mean is, does the answer to the article's question change if one is considering training vs. inference?
If we don’t need those weights at inference time, why do the computation to train them in the first place?
To go back to your ax+b example, imagine instead you are fitting a much higher dimensional model, but you don't know how high. ax^n+bx^(n-1) ... where n might be in the millions, or hundreds of millions, or?? So we know if we make the model high enough order (e.g n-1 training points will give "perfect") it will overfit, so we throw some regularization and a bit of handwavy tuning and we end up with a model of say n=7213472123 and a set of a,b .. which behaves pretty well, but from it's behavior we suspect most of them dont' matter. And maybe should be <= 2million, or whatever.
So, a few obvious questions - one is can we find a way to throw out most of the a,b,c ... to get just the core, i.e. if we throw away all |k| <= 0.00001 does it change anything (for inference). A very different question is could we decide that ahead of time (during training). A different class of question looks more like "could we have figured this out from the data".
It's a lot harder to reason about the latter questions, because the former one is empirical: After training, this one doesn't seem to do anything. Ahead of time, how do you know? This has interesting offshoots, like how stable is the distribution of the parts that matter, etc.
Yes. There's the so called "lottery ticket hypothesis". Essentially the idea is that large models start with many randomly initialized subnetworks ("lottery tickets"0 and that training finds which ones work best. Then it's only natural that during inference we can prune all the "losing tickets" away, even though we need them during training.
It's kind of an open question how large this effect is though. As the article mentions, if you can prune a lot away, this could also just mean that the network isn't optimally trained.
It strikes me as an interesting open question since if it is the case that you need big networks for training but can use significantly smaller "pruned" networks for inference there are many, many reasons why that might be true. Determining which of the possible reasons is the actual reason may be a key in understanding how LLMs work.
This is saying that you don’t need the entire model to make good predictions for specific subsets of tasks. You can literally remove a large part of the model and it will do fine. Which is not very controversial. The model, after being trained, is a large collection of interacting nodes. When this is talking about dropping chunks of the model it means dropping nodes after training to make predictions. The advantage primarily being that smaller models are cheaper and faster to run or modify with further training.
You know that meme about how you only use 10% of your brain at a time? Well, yeah, but the idiot movies that suggest using 100% of your brain would make you impossibly smarter are not correct. 90% of your brain just isn’t relevant. More brain / model is not better than the relevant subset alone.
The important question to be asking is whether you can remove large chunks of the model without hurting its ability to generally to do well on whatever you ask it.
As a very crude example, imagine you trained a simple model to predict rainfall using a weather monitor and the number of farts you did last week. The model will probably learn that the monitor is a useful and the farts are irrelevant. If this were as simple as a linear regression, you could just remove the farts coefficient from the equation and the model would come out to the same outcomes. Neural nets are not so easily observed but it’s still just dropping the irrelevant bits to whatever you’re trying to do.
Training = optimising model parameters to ‘learn’ from data.
Inference = asking the model to make a prediction, usually assuming the model is already trained.
Instead of inference, you could say running/querying the model.
So why do the larger models perform so much better…?
Just like quantization!
So the question is, perhaps: Why are big models required for compute-efficient training?
Not being facetious, I don't know the answer, but that's my best guess
More parameters makes it easier to find solutions with low-energy.
Suppose we have a product of two variables z = x * y. And now suppose that the 'correct' product is z=2, and we're learning x and y. A very good analytical solution is x=1, y=2 (or vice versa) allowing us to eliminate either x or y from our learning problem. The total energy of (x, y) in this case is 1*2 + 2*2 = 5.
However, another solution is x = y = sqrt(2), which has energy 2: this solution is much closer to the origin. The extra variable means that we have a /surface/ of solutions instead of a unique solution, so we can hone in on ones that are easier to get to using our optimizer.
As you add more variables, you can find lower and lower energy solutions.
Consider that we initialize neural networks 'near' zero, and then walk with gradient descent in some direction towards a solution. Then adding lots of extra variables - wiggle room - makes it much easier to find a solution within walking distance of the (noisy) origin.
Would be interesting to try some explicit regularization. But unfortunately you need a million bucks to an experiment on LLMs. :/
But I also disagree with takeaway.
assuming the models are identical except one is bigger then the bigger model is better because 70% of a bigger number is larger than 70% of a smaller number.
Now if you train a smaller model much longer than the bigger model (more tokens) then you are reducing the level of "under-trainedness" to some degree. at some point, you may have a smaller model that is better than that larger model.
70% of a bigger number may be larger than 70% of a smaller number but no guarantee 70% of a bigger number is larger than say 90% of a smaller number and so on.