Any time we find more efficiency, we can trade it for more quality by doing more compute. We'll always use as much compute as we can afford, until we stop getting quality gains that are worth the added cost.
Any time we find more efficiency, we can trade it for more quality by doing more compute. We'll always use as much compute as we can afford, until we stop getting quality gains that are worth the added cost.
I don't much care for all the "oh but the energy usage" claims in most tech things: it's all electricity, and it's all fungible. It usually seems to roll out as a proxy for "I don't like this thing".
Like even with cryptocurrency, there were a lot of people mistaking the issue of scalability - namely that "as a store of value" crypto would consume incredible amounts of other resources (and a lot of people got stuck trying to figure out how somehow "a hash" could be reclaimed for useful resources) to do less then alternatives, with "the energy usage itself is the problem".
Finding optimizations for LLMs is good because it means we can build cheaper LLMs, which means we can build larger LLMs then we otherwise could for some given constraint, which means we can miniaturize (or in this case specialize) more capable hardware. The thing which really matter is, can the energy usage be meaningfully limited to a sensible scaling factor given the capability that makes them useful?
Because environmentally, I can install solar panels to do zero-carbon training (and if LLMs are as valuable as they're currently being priced, this is a no-brainer - if people aren't lying about solar being "cheaper then fossil fuels").
Inferring is much cheaper and arguably provides quite a lot of value (even though I also think it is overhyped), for very little energy consumption, probably more is lost due to inefficiency for any physical product.
The 1M GPUs that meta purchased running at full (or less probably) load 24/7 is more than the energy of a single household.
The energy cost is in training, not inference.
But for the second:
Llama 1, 2, and 3 all have different architectures and needed to be trained from scratch. Llama 1 was released February 2023.
Same training story for openAI’s Sora, dalle, and 4o. All of mistral’s models Mamba, Kan, and Each version of rwkv (they’re on 6 now)
Not that this list is a result of survivor bias. It’s only looking at their published models too. Not the probably 1000s of individual training experiments that go into producing each model.
Like, if a couple of millions of people can use chatgpt in the manner they do today, would it matter if a house’s yearly energy budget was used up for that? Or 10?
To be fair, technology-wise they mostly solved this problem via proof-of-stake.
From an individual point of view you still expense enormous resources as a miner / validator in a proof-of-stake system. It's just that now the resources come in the form of lost opportunity costs for your staked tokens (eg staked Ethereum).
But from aggregated perspective of society, staked Ethereum is essentially free.
That has some parallels to how acquiring regular money, like USD, is something individuals spend a lot of effort on. But for the whole of society, printing USD is essentially free.
> Because environmentally, I can install solar panels to do zero-carbon training (and if LLMs are as valuable as they're currently being priced, this is a no-brainer - if people aren't lying about solar being "cheaper then fossil fuels").
There's still opportunity costs for that energy. Unless you have truly stranded electricity that couldn't be used for anything else.
> Finding optimizations for LLMs is good because it means we can build cheaper LLMs, which means we can build larger LLMs then we otherwise could for some given constraint, which means we can miniaturize (or in this case specialize) more capable hardware. The thing which really matter is, can the energy usage be meaningfully limited to a sensible scaling factor given the capability that makes them useful?
I agree with that paragraph. It's all about trade-offs. If we can shift the efficiency frontier, that's good. Then people can decide whether they want cheaper models at the same performance, or pay the same energy-price for better models, or a combination thereof. Or pay more energy for even better model
If you train on a cluster that costs >$1M/day to operate, the wait time is likely to be a smaller concern than the financial cost, unless you're REALLY in a hurry to beat some competitor.