Researchers run high-performing LLM on the energy needed to power a lightbulb
news.ucsc.edu
news.ucsc.edu
Code: https://github.com/ridgerchu/matmulfreellm
---
Like others before them, the authors train LLMs using parameters consisting of ternary digits, or trits, with values in {-1, 0, 1}.
What's new is that the authors then build a custom hardware solution on an FPGA and run billion-parameter LLMs consuming only 13W, moving LLM inference closer to brain-like efficiency.
Sure, it's on an FPGA, and it's only a lab experiment, but we're talking about an early proof of concept, not a commercial product.
As far as I know, this is the first energy-efficient hardware implementation of tritwise LLMs. That seems like a pretty big deal to me.
> Although they reduced the number of operations, the researchers were able to maintain the performance of the neural network by introducing time-based computation in the training of the model. This enables the network to have a “memory” of the important information it processes, enhancing performance. This technique paid off — the researchers compared their model to Meta’s state-of-the-art algorithm called Llama, and were able to achieve the same performance, even at a scale of billions of model parameters.
I disagree. The authors aren't conveniently omitting anything. They show all details in a comparison against LLama models.
Moreover, all evidence I've seen so far suggests that tritwise models can scale up to state-of-the-art sizes.
---
PS. I'm talking about the paper, not the fluffy press release.
Thanks for pointing it out!
PS. I added a PS to my comment above.
Yeah but the brain does more than predictive text.
According to the researchers,
>all we had to do was fundamentally change how neural networks work,
Compared to what? I wouldn't defend LLMs as "worth their electricity" quite yet, and they are definitely less efficient than a lot of other software, but I'd still like to see how this compares to gaming consoles, or email servers, the advertising industry hosting costs, cryptocurrency, and so on. Just doesn't seem worth pointing out the carbon footprint of AI just yet.
Of course it does. It's not like AI replaced anything you mentioned. Its carbon footprint comes on top of it.
The benefit is secondary if the end result just means more carbon dioxide.
I mean it's true it hasn't replaced anything the OP mentioned, but it has definitely replaced parts of the compute that I would normally use for e.g. searching.
Instead of google's server people are accessing OpenAi servers, which wastes even more power. How is that any better?
[1] https://www.iea.org/energy-system/buildings/data-centres-and...
Any article citing the power usage without calculating it in terms of users of queries is just trying to push an agenda by omitting how many people are using it.
Overall energy use is an important metric regardless of energy per task.
Airplanes are WAY more efficient per passenger than they were in the past, but it's still valid to express concern over the energy usage and pollutants of air travel with so many more routes being flown.
"Push an agenda"?
If they inserted a couple paragraphs saying "we estimate about 200M users a day, etc. etc.", would that add or detract from the article?
It's not relevant how they derived the number when you're reading, you only need an order of magnitude estimate, rest is distraction.
Active user count is not necessarily correlated to worthwhile consumption of resources.
This is not a fair claim to make. Is a milling machine less efficient than a clock?
It does a different thing, it's not really comparable.
I do think if we look at translation tasks, grammar correction, information look up, etc. it adds a competitive convenience factor but I can’t say that running state of the art GPUs at very high wattages for up to a minute to do what specialized software can enable you to do by running for some milliseconds on much lower wattages isn’t less efficient. I’m referring to running multiple Google searches yourself to answer a question, or using a more traditional translation service, spell checker, and so on.
What makes you confident it gave you an accurate answer?
The actual paper is here: https://arxiv.org/abs/2406.02528
The key part from the summary:
> To properly quantify the efficiency of our architecture, we build a custom hardware solution on an FPGA which exploits lightweight operations beyond what GPUs are capable of. We processed billion-parameter scale models at 13W beyond human readable throughput, moving LLMs closer to brain-like efficiency.
There is a lot of unnecessary obfuscation of the numbers going on in the abstract as well, which is unfortunate. Instead of quoting the numbers they call it “billion-parameter scale” and “beyond human readable throughout”.
> The 1.3B parameter model, where L = 24 and d = 2048, has a projected runtime of 42ms, and a throughput of 23.8 tokens per second.
e.g. 64 x 13.67W = 874 Watts to run a 1.3B model at 23.8 t/s... I'm pretty sure my phone can do way better than that! Even half that power given their assertions in the table are still overpowered for such a small model.
That's a really, _really_ big difference in memory usage and since this scales sub-linear (300M param model uses 0.21GB, 13B model uses 4.19B) a 70B model would fit on an RTX 4090. I think currently people often run 34B Models with 4bit quants on that so I would like to see some larger models trained on more tokens with this approach.
Also their 2.7B Model took 173hours on 8 NVIDIA H100 GPUs and that also seems to roughly scale linearly with the parameter size, so a company with access to a small cluster of those DGX pods (say 8) could train such a model in about 30 days - though the 100B token training set might be lackluster for SotA but maybe someone else could chime in on that.
If anyone can offer insight that would be greatly appreciated
I had expected them to make the title more clickbaity, but that number is about right for a modern lightbulb.
However, there are non-SI units that are somewhat commonplace. Horsepower, foot-pounds per second, or BTU per second aren't unheard of.
Once you do that, a dot product is just addition/subtraction. Matrix multiplications are just dot products, so you've removed multiplication.
Then they built custom hardware that presumably only does that operation and doesn't use much electricity
Since this is just "more aggressive quantization" it's not too surprising that it reaches similar performance to other quantized models.
The network should structurally look the same as any other llm, it's still using transformers etc (afaiu)
My professor at the time was at his last leg in impact, when ANNs were looked down on right before someone had the bright idea of using video cards.
I hope he’s doing well/retired on a high note.
The FPGA trinary implementation is also really interesting!
Discussion a few weeks ago: https://news.ycombinator.com/item?id=40620955