“Imprecise” language models are smaller, speedier, and nearly as accurate
spectrum.ieee.org
spectrum.ieee.org
My intuition is that a model which is undertrained suffers less from quantization, because the training process has not utilized each weight to its full potential. One of the key findings with llama, and why it punches above its weight for its size, is that they trained it for longer on a much larger dataset then was "optimal" according to the literature up to that point.
Putting two and two together, it seems that:
small model, lots of data, long training > large model + quantization
That basically, quantization is a lossy shortcut to the tail of training long. Amount and quality of data is, as always, the most important part about all of this.
For me it completely replaced strong models such as Mixtral-8x7B and DeepSeek-Coder-Instruct-33B.
1. https://www.reddit.com/r/LocalLLaMA/comments/1cst400/result_...
What, specifically, are you asking of these LLMs? "creative tasks" can be anything from programming to cooking recipes, so a tiny bit more specificality would be appreciated :)
I am too dumb for all of this ML stuff. Can you explain what exactly that means & why it's surprising?
More training leads to better compression. Trimming precision off the integers reveals that there was more compression that could be had.
Quantize each block or vector as a unit, bound to some metric instead of simply quantizing the whole thing all at once.
It's an interesting topic https://old.reddit.com/r/LocalLLaMA/comments/1ba55rj/overvie...
I did a rough calculation, and as the precision of the scalar weights decreases, the information content of the specific network permutation becomes a much higher percentage of its overall size.
Depending on the particular neural network architecture, there may be other symmetries beside the symmetric group that also represent compressible redundancies.
We’ve made impressive progress, getting these models to be quite accurate. But pushing from 90% to 99.9999999% accuracy? That takes an insane amount of data and computing power. It's like needing exponentially more energy as you get closer to light speed.
And just like we can’t actually reach the speed of light, there might be a practical limit to how accurate LLMs can get. Language is incredibly complex and full of ambiguities. The closer we aim for perfection, the harder it becomes. Each tiny improvement requires significantly more resources, and the gains become marginal.
To get LLMs to near-perfect accuracy, we'd need an infinite amount of data and computing power, which isn't feasible. So while LLMs are amazing and have come a long way, getting them to be nearly perfect is probably impossible—like reaching the speed of light.
Regardless I hope to appreciate the progress we've made but also be realistic about the challenges ahead. What do you think? Is this a fair analogy?
Beyond that, I'm not entirely sure what a "perfect" LLM would even be defined as.
I'm struggling to think of any medium that has ever reached 100% accuracy, so to target that for an ML algorithm seems foolhardy
How close to 100% is close enough?
Arithmetic and logic are close enough that cosmic rays are the limiting factor, and "computer" used to be a profession.
Synthetic data are the answers. For example see Tiny Stories dataset (https://arxiv.org/abs/2305.07759).
> Now ask the best LLM trained on "all the data" to translate some fragment of some isolate language not in its training set and not very related to existing languages.
If you give them the dictionary and grammar book as in-context instructions, it can do pretty well.
“Gemini v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning.”
Synthetic data might be the answer if you're fine with any data, but I haven't came across many synthetic datasets that are of high quality, and if you want high quality output from a LLM, I'm not sure Tiny Stories et al can provide that.
Here is just one example from Tiny Stories (https://huggingface.co/datasets/roneneldan/TinyStories/viewe...):
> Once, there was a girl who wanted to write a story. She thought and thought about what she could write about. She felt it was too boring to just write about trees and flowers. Suddenly, an idea came to her. She decided to write about her waist. She started to write about how her waist was round, and how it jiggled when she danced. Her story was so fun and exciting! She wrote about how she liked to put a belt around her waist and how it made her feel smarter. She even wrote a rhyme about her waist: "My waist is round and jiggly, And when I dance, it's so wiggly." The girl was so proud of the story she wrote. She was no longer bored - writing about her waist was much more fun!
Hardly high quality "story", and an LLM training on data like that won't have high quality output no matter how much you train it.
Edit: Another example from Tiny Stories, just because how fun they end up being:
> One day, a little boy named Jack was playing in his room. He decided to go and sit on his favourite chest. When he sat down, he noticed something unusual. The chest smelled smelly! Jack had never noticed a smelly smell before and he couldn't work out what it was. Jack's Mum heard him say 'That chest smells smelly', so she came into his room to see what was happening. When she saw the chest, she knew what was wrong. Jack's little puppy had been using the chest as a bed! His Mum scooped the naughty puppy up in her arms and took him outside. When the puppy was outside, the smelly smell went away. Jack was so relieved! He sat back down on the chest, and said 'Ahhh, much better!'
Do people really expect to be able to train on this and get high quality output? "Garbage in, garbage out", or however that goes...
>This raises the question of whether the emergence of the ability to produce coherent English text only occurs at larger scales (with hundreds of millions of parameters or more) and complex architectures (with many layers of global attention).
>In this work, we introduce TinyStories, a synthetic dataset of short stories that only contain words that a typical 3 to 4-year-olds usually understand, generated by GPT-3.5 and GPT-4. We show that TinyStories can be used to train and evaluate LMs that are much smaller than the state-of-the-art models (below 10 million total parameters), or have much simpler architectures (with only one transformer block), yet still produce fluent and consistent stories with several paragraphs that are diverse and have almost perfect grammar, and demonstrate reasoning capabilities.
The point of TinyStories isn't to serve as an example of a sophisticated model, but rather to show that the emergent ability of producing coherent language can happen at smaller scales, and from a synthetic data set, no less. TinyStories is essentially the language model equivalent of a young child, and it's producing coherent language -- it's not producing grammatically correct nonsense like the famous "colorless green ideas sleep furiously" phrase from Chomsky.
>but I haven't came across many synthetic datasets that are of high quality
I'm not really sure what your personal experience has to do with the viability of synthetic data; it's already been proven to be a useful resource. For example, Meta directly stated this upon the release of their Llama 3 model:
>We found that previous generations of Llama are good at identifying high-quality data, so we used Llama 2 to help build the text-quality classifiers that are powering Llama 3. We also leveraged synthetic data to train in areas such as coding, reasoning, and long context. For example, we used synthetic data to create longer documents to train on.
https://ai.meta.com/blog/meta-llama-3-meta-ai-responsibility...
The problem is not the amount of data, it's the quality of the data, full stop. Beyond that, there's something called the "No Free Lunch Theorem" that says that a fixed parameter model can't be good at everything, so trying to make a model smarter at one thing is going to make it dumber at another thing.
We'd be much better off training smaller models for specific domains and training an agent that can use tools deepmind style.
My understanding is NFL only applies if the target function is chosen from a uniform distribution of all possible functions — i.e. the "everything" that NFL says you can't predict is more like "given this sequence from a PRNG (but we're not telling you which PRNG), infer the seed and the function" and less like "learn all the things a human could learn if only they had the time".
At some point we'll discover a mew algorithm/architecture that can actually continuously learn from its environment with limited information and still produce amazing results like us.
But if you are searching the Internet, you'll find multiple answers — probably contradictory — and the next step is to ask the model to judge among them. Now you want all the intelligence you can muster. Unless you really trust the search engine, in which case yeah a small model seems great.
Really though I think this line of thinking is better to revisit in 5 or so years. LLM's are still very new, and seemingly everyday new optimizations and strategies are being found. Let's at least hit a plateau before assessing limitations.
You will never be able to trust a LLM more than a human expert. Because human expert will use the best available tools (for example, LLMs), will understand "what the client wants" and will put the data in the right context. At best human expert and LLM will be indistinguishable, but I really doubt it. And I think it will take a long time.
At least it's my opinion, we'll see what happens.
You're not wrong, but when that happens, does it still count as "a human expert" doing it? A chess grandmaster is capable of using Stockfish, but it's not their victory when they do.
High precision models are very expensive to run: they require expensive hardware and lots of energy.
Isn't it possible, or even likely, that low-precision models would be good enough for many tasks?
In that perspective I'm totally fine with using a 5 MB LLM like this one : https://neuml.hashnode.dev/train-a-language-model-from-scrat...
Even if there was zero efficiency gain in ternary weights, large models should probably be trained on a network of precision limited weights from here on out given the research so far.
I suspect it relates to each weight relating to multiple 'features' in the network. The greater the precision, the more room it gives for competing features to compromise on node values that aren't best for either feature instead of reorganizing the feature mapping to avoid conflicts.
For instance, one could extend dropout regularization to several levels, where each weight could have random chances to include the most significant 2-16 bits part of the time (and still 0 part of the time), and where the impact on the gradient of having fewer bits could be used to tune the ideal number of bits for each weight.
Then one could add L1 regularization for the total number of bits used to squeeze the total down down to whatever size one aims for.
And by "the comments" I'm assuming you mean the top comment [0] (and maybe its replies)? The rest don't really come off as negative at all, and you quoted directly from that top comment.
FWIW, I don't find that comment either negative or cynical. It starts out with that sentence you quoted, but it goes on to make a very interesting point about quantization most likely working best for models which are undertrained—models which store less information than their architecture size would suggest. That's a very valid point that I found insightful and interesting, not cynical.
Combine that with what seems like every chip adding a neural engine and it feels like we are in the early days of high performance graphics again. Right now we are in the unreal engine voodoo era, where graphics cards/neural engines are expensive/rare. But give it a few generations and soon we can assume that even standard computers will have pretty decent NPUs and developers will be able to rely on that for models
Is this the level of performance people are relying upon? While I've always been impressed with the technology itself, it's only starting with GPT 4 that I think it approaches adequate performance.
The way I rationalize it is that using 3.5-turbo is like programming on an 8-bit computer with Kilobytes of Ram, and gpt-4o is like programming on a 64bit computer with a 4080 ti and 32gb of ram. If I can make things work on the 8-bit system, they will work nicely on the more powerful system.
What's an acceptable rate? Is it 99%? 99.9%?
The closer it gets to 99.999% good answers, the more damaging the wrong ones become. Because people have been trained, too. Trained to trust the answers, which makes them lazy and vulnerable to lies.
I don't think that this is just a nit, it implies a real mismatch between how these models work, and how modern computing hardware works. People have built ternary computers in the past, but to the best of my knowledge nobody's made a serious attempt at it in half a century.
You can always use two bits to store a ternary value, of course. But then it wouldn't be a 1-bit LLM, now, would it? And that doubling in compute resources required would make a tangible difference in how one has to think about potential efficiency improvements. Also, ternary tritwise logic implemented on binary hardware is unlikely to be anywhere near as efficient as binary bitwise logic on hardware that was specifically engineered for the purpose. This leaves me thinking that the research teams involved continuing to refer to these as 1-bit LLMs must be interpreted as knowingly engaging in dishonest marketing tactics.
BitNet is (I think) actually one bit per parameter. BitNet 1.58b isn't, but then to be fair isn't describing itself as a 1 bit llm. I'm less sure but it seems OneBit is 1 bit. One of the processes mentioned in the article is a mix of one and two bits for the weights.
For those unaware, the MsoTriState is fairly self-explanatory: it is a tri-state boolean type with five possible values.
If you have a trillion parameter 8-bit fp network or a trillion parameter 1.5-bit ternary network, based on the scaling in Microsoft's paper the latter will actually perform better.
A lot of the current thinking is that the nodes themselves act as superpositions for a virtualized network in a multidimensional vector space, so precision is fairly arbitrary for the base nodes and it may be that constraining the individual node values actually allows for a less fuzzy virtualized network by the end of the training.
You could still have a very precise 'calculator' feature in the virtual space no matter the underlying parameter precision, and because each parameter is being informed by overlapping virtual features, may even have less unexpected errors and issues with lower precision nodes.
LLMs are sufficiently good hammers that people see everything as a nail, then talk about how bad they are at driving screws.
I’m not a theoretician nor scientist, but train models on math and logical reasoning. Some very simple word problems contains tons of context, like family trees, implicit information, etc. and the process of step by step reasoning, not just the final answer, requires even more natural language processing.
Tangent but... programming languages aren't designed for computers. Computers are perfectly happy with assembly or even binary. Programming languages are designed for humans, not just so we can see what others have done, but so that we can understand what we ourselves have done. We give the variable a name because we can't remember 0x0010101; but the computer remembers them both just fine.
The "nearly as accurate" is only on their contrived benchmarks. I've never met a quantized model that actually behaved "98%" as good as the unquantized model, and I do LLM work daily and have since well before the ChatGPT era.
Honestly it feels like the bulk of the industry is acting out one big LARP where they're always right around the corner from developing AGI and doing something big and amazing and....it just never materializes. Obnoxious AI agents, worthless AI-generated websites crowding out useful resources, unreliable AI search results.
The AI industry has done very well for itself selling hype. Now it needs actual good products.
I don't disagree that there is a ton of hype, a lot of it unwarranted, and I cringe when I see tons of companies trying to "throw AI against the wall and see if it sticks" (I personally nominate LinkedIn's AI blurbs on their feed as "most useless and annoying use of AI"). But still, I'm blown away with how much value I get from AI. It makes me a bit sad that so many of us have become so jaded that they see it as "one big LARP".
I can't trust ChatGPT.
If I am searching for something I don't know the answer to and don't have the luxury of trial-and-error for the information I'm given, I can't rely on an unreliable agent like ChatGPT (or literally any LLM for that matter).
ChatGPT could be giving me a correct answer. Or it could be blowing smoke up my ass.
I don't know which it is when I'm seeing an answer from ChatGPT!
And that's the problem.
At least not my stuff where it'd be quicker to ask an AI agent for help with.
If you’re going to verify (and you should), might as well skip the asking step.
I asked chatgpt for recommended food pairings for a soup I made today. It did a great job.
Chatgpt also helped me debug a home automation issue I'd been having for over a year with some smart lights.
I find uses for chatgpt every day.
She said the executives told the board they "didn't think he was the right person to lead the company to AGI…”
The way I read it, it sort of sounded like you could substitute “AGI” with “Shangri-La”.
It’s always going to be just down the road, but they are sort of emotionally convinced that is where they are headed.
It's been less than two years since ChatGPT released.
I won't go over the entire table in detail, but PIQA, BoolQ, HellaSwag, and WinoGrande should in the mid-to-high 70s for LLaMa2-7B. They drop that to 58, 62, 32, and 51. There are 700M parameter models that perform much better.
What they should have reported is effective number of parameters. Does LLaMa2-7B with their quantization method outperform a model that has amount of computation but uses that compute with say.. 16-bit quantization? If the answer is no, and it seems like it very clearly is, then the method is wholly worthless. Just use a smaller model to begin with.
The BitNet paper is better. But for some reason they only consider very small models. It's the obvious question and in their FAQ they don't provide a solid answer to it. Despite having all of the compute resources of MS. They could have easily run this experiment in the past year; I'm suspicious.
There's also no practical reason. Training quantized networks is often harder! This is why people quantize after the fact or do distillation.
Nor is there any reason why we won't find some projection of weights onto the BitNet manifold.
If it was published by academics I'd believe the cost argument.
This was published by MS. They can run this experiment trivially. I have friends at MS with access to enough compute to do it in days.
Either the authors ran it and saw it doesn't work or they're playing with us. Not a good look. The reviewers shouldn't have accepted the paper in this state without an explanation of why the authors can't do this.
This is the question that determines if this work matters or is useless. Publishing before knowing that isn't responsible on anyone's part.
* https://huggingface.co/1bitLLM/bitnet_b1_58-large
* https://github.com/Oxen-AI/BitNet-1.58-Instruct
* https://github.com/nkotak/1.58BitNet
See some followups that has some training advice: https://github.com/microsoft/unilm/tree/master/bitnet
If anyone is running it here in a production setting, please post and prove me wrong.
So why do people keep focusing on "1-bit" when the whole reason 1-trit models are so successful in the first place just might have everything to do with the ternary weights and the symmetry they're able to encode?
I'm not sold on the suggestion that "imprecise language models" in general are "nearly as accurate" when it might actually be that ternary weights are more precise in a totally difference sense: in that they just might be capturing a minimal representation of the very symmetries responsible for making all those parameters effective in the first place.
And the whole result is also conditional on the optimization/training process used. Which is an area where we have no reason to think that we are optimal... So we can do studies with practical results (given sufficient money), but we are far from being able to identify the actual maximums available.
I've seen lots of headlines like "1-bit quantization" (including this one and https://arxiv.org/abs/2310.16795). What I've found in this space is that the headlines can often be intentionally misleading about what is actually achieved. If you read closer the abstract of this paper, it claims 8.41 perplexity on LLaMA2-70B at 1 bit, which is a HUGE decrease from 3.120 perplexity in FP16 and they will never mention that in the headline. Even LLaMA2-7B at INT8 achieves 5.677 perplexity with half of storage place (better with LESS space and LESS training). Some claim 1.58-bit quantization (each weight is either -1, 0 or 1) but in practice require very small group sizes, which means one or two extra FP16 numbers for every 64 weights and that means another 0.5 bit, so it's actually 2-bit quantization. And every quantization algorithm can claim they make language models smaller, speedier, and use less energy, so there's nothing special about these.
Here are the key metrics that I suggest checking when comparing quantization schemes:
* Perplexity. Note that it also depends on the dataset (either WikiText2 or C4, WikiText2 numbers are usually lower than C4) and context size (1024, 2048 or 4096, higher context sizes usually means less perplexity). Dataset and context size must match to make a meaningful comparison.
* Quantization bits. Many algorithms claiming 2-bit or 1-bit quantization has lots of extra parameters elsewhere, such as grouping. Download the quantized version of the model, check its file size, multiply by 8 and divide by the number of parameters. That gets you the ACTUAL quantization bits.
* Performance. Weights may need to be dequantized during inference which could introduce overhead. Some libraries have custom matmul kernel for dequantization that achieves performance close to FP16, others can be slower. Check its generation speed and inference speed.
Newer architectures such as Ampere contains INT8 cores which may make quantized version even faster than FP16, I haven't tried out yet.
There is also a lot of misleading comparisons in this space. Some methods like GPTQ only provide vector-matrix multiplication kernels, which means only a single token can be generated and batched inference (which is needed for generating initial KV cache, or for serving multiple users) can be much slower. If an algorithm claims a 3x speedup for something, check if they refer to single stream latency or multiple stream throughput. Some of that speedup comes from running a model on 2 cards instead of 5 cards without specifying if the cards have NVLink configured (you shouldn't run inference on multiple cards without NVLink, or you should expecet huge slowdown simply because of using 5 cards).
* Base model. Pick a STRONG base model like Llama-2-7b or Llama-3-8b etc. Not an undertrained model like SwitchTransformer etc which may have lots of redundant parameters in itself.
My personal favourite remains QuIP# (https://github.com/Cornell-RelaxML/quip-sharp). It lacks in the "performance" part as its matrix multiplication performance isn't on par yet but there is room for improvement, and it wins every other metric. And sad news: it's very likely we won't have practical 1-bit LLMs, never ever. We are reaching the end game between 2.5~4 bits. By "practical" I mean it should beat 3-bit LLMs with 3x less parameters or 2-bit LLMs with half as many parameters. There is a Shannon limit to quantization whatever methods you use.
While I'd love to see it scaled up to at least ~50B models, it looks like limited weight precision might actually offer improved network optimization over unconstrained weights for pretraining.
Do you think that work is misrepresenting the gains, or that QAT is a different beast where quantization isn't as much a tradeoff as a potential net gain across the board?
> Mixed precision training. While the weights and the activations are quantized to low precision, the gradients and the optimizer states are stored in high precision to ensure training stability and accuracy. Following the previous work [LSL+21], we maintain a latent weight in a high-precision format for the learnable parameters to accumulate the parameter updates. The latent weights are binarized on the fly during the forward pass and never used for the inference process.
In this case there are two areas to optimize for: training efficiency and inference efficiency.
If I understand correctly, it stores the weights, gradients and second-moment estimates in FP32 like every other mixed-precision training (the Gopher paper has details on why storing them in FP32 is important), and quantized weights are used in forward pass. What I'm not sure is whether latent weights are used in backward pass, and my instinction is that the "Straight-through estimator" requires high-precision latent weights so they may still be needed. Training FLOPS can be roughly estimated as 6 FLOP per parameter per token, where 2 is forward pass, 2 is gradient computation and 2 is gradient accumulation (see https://medium.com/@dzmitrybahdanau/the-flops-calculus-of-la...). If only forward pass is quantized, this means only 1/3 of all FLOPS are optimized (and even then it has to be accumulated in FP32). So I'm skeptical of the gains in training efficiency here, and I can't find the numbers (how much energy or how much time is used for training, compared to regular FP16 mixed precision training? The papers boast inference energy savings which makes me even more skeptical of training energy savings)
For quantization efficiency, while QAT can certainly avoid the quantization step, PTQ methods are very cheap (usually <24 hours on RTX 4090 for Llama-2-70b) so I consider the cost of the quantization step negligible. There is not much difference in inference efficiency gains as PTQ and QAT can quantize to the same format. For final accuracy, unfortunately there is a lack of comparison between QAT and PTQ of fp16 models, and PTQ has the advantage of not requiring access to the original dataset, so I think it's very hard to make a fair comparison here but it's also likely the only area where QAT has actual gains compared to best PTQ methods.
What I don't think that tells us anything about is directly trained 1-bit LMs versus 3-bit LMs, because in that case there's no compression step to introduce quantisation artefacts. There might be an analogous training data size argument but it's not clear to me that there needs to be: a 3X parameter 1-bit LLM and a 1X parameter 3-bit LLM ought to be equivalent in terms of their information capacity.
I would want to continue the statement that we are still early innings on renewable energy -- and let's keep deploying it rapidly to manage increased compute demand.
Any time we find more efficiency, we can trade it for more quality by doing more compute. We'll always use as much compute as we can afford, until we stop getting quality gains that are worth the added cost.
I don't much care for all the "oh but the energy usage" claims in most tech things: it's all electricity, and it's all fungible. It usually seems to roll out as a proxy for "I don't like this thing".
Like even with cryptocurrency, there were a lot of people mistaking the issue of scalability - namely that "as a store of value" crypto would consume incredible amounts of other resources (and a lot of people got stuck trying to figure out how somehow "a hash" could be reclaimed for useful resources) to do less then alternatives, with "the energy usage itself is the problem".
Finding optimizations for LLMs is good because it means we can build cheaper LLMs, which means we can build larger LLMs then we otherwise could for some given constraint, which means we can miniaturize (or in this case specialize) more capable hardware. The thing which really matter is, can the energy usage be meaningfully limited to a sensible scaling factor given the capability that makes them useful?
Because environmentally, I can install solar panels to do zero-carbon training (and if LLMs are as valuable as they're currently being priced, this is a no-brainer - if people aren't lying about solar being "cheaper then fossil fuels").
Inferring is much cheaper and arguably provides quite a lot of value (even though I also think it is overhyped), for very little energy consumption, probably more is lost due to inefficiency for any physical product.
The 1M GPUs that meta purchased running at full (or less probably) load 24/7 is more than the energy of a single household.
The energy cost is in training, not inference.
But for the second:
Llama 1, 2, and 3 all have different architectures and needed to be trained from scratch. Llama 1 was released February 2023.
Same training story for openAI’s Sora, dalle, and 4o. All of mistral’s models Mamba, Kan, and Each version of rwkv (they’re on 6 now)
Not that this list is a result of survivor bias. It’s only looking at their published models too. Not the probably 1000s of individual training experiments that go into producing each model.
Like, if a couple of millions of people can use chatgpt in the manner they do today, would it matter if a house’s yearly energy budget was used up for that? Or 10?
To be fair, technology-wise they mostly solved this problem via proof-of-stake.
From an individual point of view you still expense enormous resources as a miner / validator in a proof-of-stake system. It's just that now the resources come in the form of lost opportunity costs for your staked tokens (eg staked Ethereum).
But from aggregated perspective of society, staked Ethereum is essentially free.
That has some parallels to how acquiring regular money, like USD, is something individuals spend a lot of effort on. But for the whole of society, printing USD is essentially free.
> Because environmentally, I can install solar panels to do zero-carbon training (and if LLMs are as valuable as they're currently being priced, this is a no-brainer - if people aren't lying about solar being "cheaper then fossil fuels").
There's still opportunity costs for that energy. Unless you have truly stranded electricity that couldn't be used for anything else.
> Finding optimizations for LLMs is good because it means we can build cheaper LLMs, which means we can build larger LLMs then we otherwise could for some given constraint, which means we can miniaturize (or in this case specialize) more capable hardware. The thing which really matter is, can the energy usage be meaningfully limited to a sensible scaling factor given the capability that makes them useful?
I agree with that paragraph. It's all about trade-offs. If we can shift the efficiency frontier, that's good. Then people can decide whether they want cheaper models at the same performance, or pay the same energy-price for better models, or a combination thereof. Or pay more energy for even better model
If you train on a cluster that costs >$1M/day to operate, the wait time is likely to be a smaller concern than the financial cost, unless you're REALLY in a hurry to beat some competitor.
https://hachyderm.io/@inthehands/112006855076082650
> You might be surprised to learn that I actually think LLMs have the potential to be not only fun but genuinely useful. “Show me some bullshit that would be typical in this context” can be a genuinely helpful question to have answered, in code and in natural language — for brainstorming, for seeing common conventions in an unfamiliar context, for having something crappy to react to.
> Alas, that does not remotely resemble how people are pitching this technology.