What happens if we remove 50 percent of Llama?
neuralmagic.com
neuralmagic.com
My intiution tells me the pre-training paradigm will shift immensely in near future because we started to understand that we don’t need all these paramaters since the subnetworks seems to be very robust preserving information in high dimensions. We keep saying curse of dimensionality but it is more like the bliss of dimensionality we keep seeing. Network redundancy still seems to be very high given BitNet is more less comparable to other LLMs.
This basically shows over 50% of the neural net is gibberish! The reason being is that the objective function simply does not include it.
Again my intiution tells me that neural scaling laws are incomplete as they are because they lack the efficiency parameter that needs to be taken into account (or simply left out due to greed of corporate).
And this is what we are seeing as “the wall”.
I am no expert in neural network theory nor in math but I would assume the laws should be something in the vicinity of this formulation/simulation:
https://colab.research.google.com/drive/1xkTMU2v1I-EHFAjoS86...
and encapsulate shannon’s channel’s capacity. I call them generalized scaling laws since it includes what it should include in the first place: entropy.
If I remember correctly, their counter-intuitive result was that big overparameterized models could learn more efficiently, and were less likely to get trapped in poor regions of the optimization space.
[This is also similar to how introducing multimodal training gives an escape hatch to get out of tricky regions.]
So with this hand-wavey argument, it might be the case that two-phase training is needed: A large overcomplete pretraining focused on assimilating all the knowledge, and a second that makes it compact. Other, that there is a hyperparameter that controls overcompleteness vs compactness and you adjust it over training.
Vs trying to fill something with just a narrow tube, you spill most of what you put in.
This is a mischaracterization of sparsity. Performance did drop, so the weights are not gibberish. Training vs pruning, you can't train into the final state, you can only prune there.
Going forward we will accumulate truly useful data at a linear growing rate. This fundamentally breaks the scaling game. If your model and compute expand exponentially but your training data only linearly, the efficiency won't be the same.
Synthetic data might help us pad up the training sets, but the most promising avenue I think is to use user-LLM chat logs. Those logs contain real world grounding and human in the loop. Millions of humans doing novel tasks. But that only scales linearly with time, as well.
No way around it - we only once had the whole internet for the first time in the training set. After that it's linear time.
I have a copy of refined web locally so I have a billion pre-chatgpt documents for my long term use.
> Also, most of them are silent. They don’t really do much. Or their activities are… You have to hit it with just the right set of stimulus.
> ... When you place these electrodes, again, within this hundred micron volume, you have 40 or so neurons. Why do you not see 40 neurons? Why do you see only a handful? What is happening there?
(Yes, I understand LLM aren't brains.)
During the learning stage we want input from every variable so that we are sure that we don't omit a variable that turns out to be essential for the calculation. However in any calculation a human does 99.9999% of variables are irrelevant (e.g. what day of the week it is, am I sleepy, etc), so of course the brain wouldn't use resources to keep connections that aren't relevant to a given function. Imagine what a liability it would be if we have had excessive direct connections from our visual processing system to the piece of our brain that controls heartrate.
The unfortunate reality is that no one truly understands how memory works. Many theories are floating around, but the fundamental components remain elusive. One thing is certain: it is quite different from backpropagation. Thankfully, our brains do not suffer from catastrophic forgetting.
The only paper I have seen claiming this studied only lightweight open-source models (<27B, mostly 2B and 8B). The also included o1 and 4o for reference, which kind of broke their hypothesis, but they just left that part out of the conclusion. Not even kidding, their graphs show o1 and 4o having strong performance in their benchmarks, but the conclusion just focuses on 2B and 7B models like gemma and qwen.
An 18% drop in accuracy (figure 8) is not insignificant. Even 4o suffered 10% loss (figure 6), and 4o isn't a small llm.
Competent performance should have near zero performance loss. The simplest benchmark merely changes things like "john had 4 apples" to "Mary had 4 oranges." Performance loss due to inconsequential tokens changing is the very definition of over-fitting.
The authors did the equivalent of "Lets design a human intelligence benchmark, and use a bunch of 12 year olds as reference points"
I will eat my hat if the authors rescind the paper in a year or so if their benchmarks show no difference on SOTA models.
No definitive answer yet, but my bet is on no.
Large models successful now have dodged recurrent architecture, which is harder to train but allows for open ended inference steps, which would allow straightforward scaling to any number of reasoning steps.
At some point, recurrent connections are going to get re-incorporated into these models.
Maybe two stage training. First stage, learn to integrate as much information as well as possible, without recurrence. As is happening now. Second training stage, embed that model in a larger iterative model, and train for variable step reasoning.
Finally, successful iterative reasoning responses can be used as further examples for the non-iterative module.
This would be similar to how we reason in steps at first, in unfamiliar areas. But quickly learn to reason with faster direct responses, as we gain familiarity.
We continually fine tune our fast mode on our own more powerful slow mode successes.
Still 5k points to go, though! :D
Those models (4o, o1-mini, preview) don't see any drop at all on those benchmarks. The only benchmark that see drops with the SOTA models is the one they add, "seemingly relevant but ultimately irrelevant information".
Humans can and do drop in performance when presented with such alterations. Are they better than LLMs in that case ? Who knows ? Because these papers don't bother testing human baselines.
- after the answer, ask it "are you sure?" (from the office tv series: "is it a stupid thing to do? if it is, don't do it") - chain of thought, step-by-step thinking - different hats (godfather style: piecetime vs. wartime consigliere): looking at the problem from different points of view (at the same time or in stages). For example, first draft: stream of consciousness answer, second iteration: critic/editor/reviewer (produces comments), third (address comments), repeat for some time - collaborative work of different experts(MoE), delegate specific tasks to specialists - [deliberate] practice with immediate feedback
Is it possible to build domain specific smaller models and merge/combine them at query/run time to give better response or performance instead of one large all knowing model that learns everything ?
Nobody understand how LLM works either. Is LLM as "early" as MoE ?
If you will excuse analogy and anthropomorphism, the human analogy of what we do and don't understand about LLMs is, I think, that we understand quantum mechanics, cell chemistry, and overall connectivity (perceptrons, activation functions, and architecture) and group psychology (general dynamics of the output), but not specifically how some belief is stored (in both humans and LLMs).
Any citation on this one?
Also note that I’m ELI5 so saying word is fine.
https://proceedings.neurips.cc/paper_files/paper/2023/file/d...
You can use a specific LLM, or a general larger LLM to do this routing.
Also, some work suggest using smaller llms to generate multiple responses and use a stronger and larger model to rank the responses (which is much more efficient than generating them)
A common practice in more formal domains is to have a portfolio of solvers and race them, allowing for the first (provably correct) solver to “win”
In less formal domains, adding/removing nodes/trees in an online manner is part of the deployment process for random forests.
Would love to know more about how they filtered the training set down here and what heuristics were involved.
I think that the models we use now are enormous for the use cases we’re using them for. Work like this and model distillation in general is fantastic and sorely needed, both to broaden price accessibility and to decrease resource usage.
I’m sure frontier models will only get bigger, but I’d be shocked if we keep using the largest models in production for almost any use case.
It is significantly more complex than it appears at first sight.
I wonder how much they'd be able to trim the recent QwQ-32b. That thing is actually good enough to be realistically useful, and runs decently well with 4-bit quantization, which makes it 16Gb large - small enough to fit into a 3090 or 4090, but that's about it. If it can be squeezed into more consumer hardware, we could see some interesting things.
If a 32B model@4bit normally requires 16 GB VRAM, at half the size, it could be run @8bit with 16 GB VRAM?
Isn't that tradeoff a great improvement? I assume the improved bit precision will more than compensate for the loss related to removal?
The other thing is that VRAM is used not just for the weights, but also for prompt processing, and this last part grows proportionally as you increase the context size. For example, for the aforementioned QwQ-32, with base model size of ~18Gb at 4-bit quantization, the full context length is 32k, and you need ~10Gb extra VRAM on top of weights if you intend to use the entirety of that context. So in practice, while 30b models fit into 24Gb (= a single RTX 3090 or 4090) at 4-bit quantization, you're going to run out of VRAM once you get past 8k context. Thus the other possibility is that VRAM saved by tricks like sparse models can be used to push that further - for many tasks, context size is the limiting factor.
I assume in your post that "30b" meant 30 billion, or in other words, 30Gp (giga-parameter).
Furthermore is 24Gb of VRAM 24 gigabits (power of 10), or 24 gibibits (power of 2)?
With VRAM, this quite obviously refers to the actual amount that high-end GPUs have, and I even specifically listed which ones I have in mind, so you can just look up their specs if you genuinely don't know the meaning in this context.
But perhaps that's just me…
LLM inference workloads are bound by the compute power, sure, but that's not insurmountable IMO. Much bigger challenge is memory. Not even the bandwidth but just a sheer amount of RAM you need to just load the LLM weights.
Specifically, even a single H100 will hardly suffice to host a mid-sized LLM such as llama3.1-70B. And H100 is ~50k.
If that memory amount requirement is there to stay, and with current LLM transformer architecture it is, then what is really left as an only option for affordable consumer HW are only the smallest and least powerful LLMs. I can't imagine having a built-in GPGPU with 80G of on-die memory. IMHO.
Massive increases in demand due to this stuff being really really useful can cause prices to go up even for existing chips (NVIDIA is basically printing money as they can sell all they can make at for as much money as the buyers can get from the investors). I have vague memories of something like this happening with RAM in the late 90s, but perhaps it was just Mac RAM because the Apple market was always its own weird oddity (the Performa 5200 I bought around then was also available in the second hand listings on one of the magazines for twice what I paid for it).
Likewise prices can go up from global trade wars, e.g. like Trump wants for profit and Biden wants specifically to limit access to compute because AI may be risky.
Likewise hot wars right where the chips are being made, say if North Korea starts fighting South Korea again, or if China goes for Taiwan.
We don't even need to go that far in the history. Crypto hype just few years ago skyrocketed the GPU prices.
Nonetheless, great info. Sounds like it might be the budget inference king!
The workstation edition of GPUs usually do the clamshell configuration so they can easily double the VRAM and ramp up the price by a couple thousand
Compounding that four times, we should get .8^4 = 40% of the accuracy for .2^4 = .16% of the size.
That’d be about 1 GB for the current largest model.
Something tells me that's a little optimistic.
The gargantuan # of parameters is what buys you the generalization properties that everyone is interested in. A very reduced model may still look & sound competent on the surface, but extensive use by domain experts would quickly highlight the cost of this.
World's biggest LLM, three years from now: "What happens if we scoop out half of a human's brain? Probably not anything significant."
https://www.cbc.ca/radio/asithappens/as-it-happens-thursday-...
Understanding this could probably make the problem easier by some factor (but not "easy" in any sense.)
I was going to write "I don't think this specifically is where we need to look", but then I remembered there's two different reasons for mind uploading.
If you want the capabilities and don't care either way about personhood of the uploads, this is exactly what you need.
If you do care about the personhood of the uploads, regardless of if you want them to have it (immortality) or not have it (a competent workforce that doesn't need good conditions), we have yet to even figure out in a rigorous testable sense what 'personhood' really means — which is why we're still arguing about the ethics of abortion and meat.
I was only thinking just about malicious de-humanising, but yes, you're right, that absolutely is a valid example.
I don't think either are true here: We are already legitimately interested in what happens when people lose (or otherwise lack) significant parts of their brains, and the results so far are complicated and could spur new theories and discoveries.
Functioning autism hardly equals low intellect. Half the people of this forum (at least) are functioning autists.
That said, the parent comment is just silly and wrong.
Make sense?
For others, ID plays no part. even at the subdiagnostic level.
I would imagine some expressions of autism are automatically called intellectual disability just because it's not understood well enough for people to effectively teach for it. Of course people will think you're intellectually disabled if you learn so significantly differently that most of the education that works well for most other people does not work nearly as well for you. That doesn't necessarily make you intellectually disabled though, it just makes you bad at doing the same thing as everyone else. Which, to be fair, is already the source of practically all of the social consequences of neurodivergence.
My particular flavor of autism seems to make me a decent programmer (I hope so, anyway). While regular education also does not entirely work for me, I did self-learning that allowed me to still keep up in school, and I can still somewhat benefit from resources made for non-autistics, just not always as much as one is "supposed" to.
I personally benefit the most from explanations of how something is implemented rather than directions to achieve certain arbitrary goals, because if I know how something is implemented, I will be able to achieve any goal with it. Nowadays, I can usually manage to figure out how something is implemented based on directions, so I can still usually learn from directions, it's just much slower / less efficient for me.
I'd imagine forms of autism that are automatically called intellectually disability might not be able to "work backwards" like this, even if they are perfectly capable of autistic logic and reasoning, just because they haven't yet developed the skill needed to extract value from directions.
I sincerely hope that further research in this field will finally reveal how to effectively help these people rather than just calling them disabled in the context of a neurotypical education.
For the rest of us, it's not always a gift. It can be (for me that's analytical thinking and technical writing). But it can also be an absolute curse.
Wow. Absolutely not. ADHD has practically ruined the majority of my life and BPD ruins many/most of my social interactions. That doesn't mean there's anything wrong with me though, it just means I'd rather die, but I prefer not to impose that view on others. I'm relatively proud-ish to be myself despite my flaws, that absolutely does not mean I don't have flaws or that they haven't caused me immeasurable pain. Go call privileged someone else.