Numbers every LLM developer should know
github.com
github.com
- Human Reading Speed (English): ~250 words per minute
- Human Speaking Speed (English): ~150 words per minute
Should be treated like the Doherty Threshold [1] for generative content.
Is it? I've noticed a huge variance in speaking speed in the US, but it tends to vary more between regions rather than individuals.
Ask GPT-4 a question and then answer it yourself. Maybe your answer will be as good or better than GPT-4's but GPT-4 writes its answer a lot faster.
I'm not sure this is accurate. From what I have seen, 8-bit quantization is usually fine, and even 4-bit is a viable tradeoff. Here are some benchmarks from TextSynth showing no significant degradation between 16 and 8 bit:
https://textsynth.com/technology.html
8-bit uses half as much memory and doubles the throughput for limited quality loss.
[1] https://blog.novelai.net/anlatan-acquires-hgx-h100-cluster-4...
https://blog.novelai.net/text-model-progress-is-going-good-8...
PS, as a joke, they should implement GPU fluint8 and get baked in non-linearity for the activation function without even using a non-linear function, https://www.youtube.com/watch?v=Ae9EKCyI1xU ("GradIEEEnt half decent: The hidden power of imprecise lines" by suckerpinch)
Nonetheless people do tend to use 16 bit huggingface models, and if you do go to 8 bits and it's wrong, you're never quite sure if it's the quant or the model.
No, 4bit quantization is the typical case.
At 4bit you can fit twice the parameters of 8bit in the same space for far better performance/perplexity/quality.
Running LLMs higher than 4bit is atypical and almost always sub-optimal (compared to running a model half the size in 8bit).
Even pretraining and finetuning in 4bit is likely to become the norm soon as fp4 becomes more well understood.
https://github.com/ggerganov/llama.cpp#quantization
4 bit has a perplexity score 0.13 or so higher.
If you are limited to X RAM and have two 16bit models of size 4X and 2X then the 4X model in 4bit will always be far superior to the 2X model in 8bit, with far lower perplexity.
Compare 13B's 4bit perplexity of 5.3607 to 7B's 8bit perplexity of 5.9069. That is over 0.54 lower perplexity for the same RAM amount by using 4bit! That is MASSIVE!
You have to wonder if running a huge model, say, 300B parameters at 2-bit quantization might be "optimal" in that it would fit into a single A100 or H100 GPU and likely outperform an 80B parameter 8-bit model...
And some other models have more crazy numbers with even more crazier outliers within them, like you might have a weight of 12.00 between long array of typical small numbers around 0.00
I've read story about attempt to quantize RWKV model into the 4/5 bits which failed short due to the presence of outlier weights.
The author told somewhere that bigger models had worse perplexity because of this.
https://github.com/saharNooby/rwkv.cpp/issues/12
For LLaMA models - yeah, different story.
You can see it in real time when you take most LLMs and compare them at different quantization levels. I can see the degradation even in the largest llama quite badly even at 8 bits.
If you have X amount of VRAM and can fit a 16bit model of size 2X in 8bit or a model of size 4X in 4bit then the 4X model in 4bit is ALWAYS superior with lower perplexity and better performance.
You LOSE performance by using a smaller model in 8bit vs a larger model in 4bit.
More detail than you probably wanted: https://huggingface.co/blog/hf-bitsandbytes-integration
Also note that for a fixed memory (RAM) size, 4bit (even int4) is always superior, resulting in lower perplexity than 8bit.
E.g. LLaMA-13B int4 is far better/lower perplexity than LLaMA-7B fp8 while using the same amount of RAM.
So in this case one would load a byte which would have 2 4b data, and then you would have a 4b ADD or MAC which would operate on them.
If you don't have them then you need to sign/zero extend or convert the smaller bit-widths to 8/16/32b whichever is available.
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
There's too many schemes right now with 4_0 and 5_1 really popular between LLM geeks.
I think that's a typo there too, the 13B model needs like 10G of memory for 4 bits, it's the 7B one that fits into 6G. Well unless you do the split thing with some layers on the CPU I guess.
MosaicML claims they trained a 7 billion parameter on 1 trillion tokens with a budget of $200k.
https://www.mosaicml.com/blog/mpt-7b
Does training cost scale linearly with model size and token count? If so, that suggests a lower bound of $600k to train the 13 billion params model. (Still roughly the same magnitude)
And these are impossible to get. We tried to get some for Anyscale, and we were told there were no on-demand available and lead time for reserved (ouchie on the price! You're talking a quarter of a million dollars a year for one machine at list) was in weeks.
Once you take the model size and hefty sweetheart deals into account, you're within 10%. Mosaic does have some nice whitebox optimizations, but nothing that radically changes the equation.
I have a suggested modification. You are mixing references in your document.
Re: '~$1 million: Cost to train a 13 billion parameter model on 1.4 trillion tokens The LLaMa paper mentions it took them 21 days to train LLaMa using 2048 GPUs A100 80GB GPUs.'
The LLaMA-13B model took 2.75 days of 2048xA100 (135,168 GPU-hours) with 1 trillion tokens. The 21 days for 1.4 trillion was for LLaMA-65B.
I would suggest using the LLaMa-13B numbers since those are the most relevant for this section, or at least modify "21 days to train LLaMa" to "21 days to train LLaMa-65B" for clarity.
i wonder when we are getting docker for llm ... a Modelfile ?
FROM "PAAMA/16b"
APPLY "MNO/DATASET"
each layer could be lora adapter like thing maybe.
maybe when AI chips are finally here.
SELECT * FROM iris.train
TO TRAIN DNNClassifier
WITH model.hidden_units = [10, 10], model.n_classes = 3, train.epoch= 10
COLUMN sepal_length, sepal_width, petal_length, petal_width
LABEL class
INTO sqlflow_models.my_dnn_model;
No idea how well it works.https://pytorch.org/tutorials/beginner/pytorch_with_examples...
Looks to me like "numbers every LLM user needs to know".
There are some unique assumptions being made in parts of the gist
> 10: Cost Ratio of OpenAI embedding to Self-Hosted embedding
> 1: Cost Ratio of Self-Hosted base vs fine-tuned model queries
I don't know how useful these numbers are if you take away the assumptions that self-hosted will work as well as API.
> 10x: Throughput improvement from batching LLM requests
I see that the write up mentions memory being a caveat to this, but it also depends on the card specs as well. Memory Bandwidth / TFLOPs offered by say 4090 is superior while having the same amount of VRAM as 3090. The caveat mentioned with token length in the gist itself makes the 10x claim not a useful rule of thumb.
In a narrow use-case of a strict look-up. This seems to exaggerate the cost difference while having completely different trade-offs.
From the phrasing around fine tuning right now it seems like it's using openai's fine tuning api to determine that cost, but it's not very clear.
Also this would be helpful for other foundation models if that doesn't already exist - how much VRAM to run Stable Diffusion v2.1 at different resolutions, running Whisper or Bark for audio, etc.
1 GPT4 token is equivalent to 50 GPT3.5 tokens.
1 token is equivalent to 0.75 words.
I'm also confused about this:
> ~$1 million: Cost to train a 13 billion parameter model on 1.4 trillion tokens
This is apparently related to the LLaMa paper, but that paper seems to cite 1.0T tokens (rather than 1.4T tokens) for the 13B model. Also, if 20 to 1 is in fact optimal for the data-to-parameter ratio, then using a 100 to 1 ratio doesn't seem like an appropriate way to arrive at a magic number for training costs. The magic number should really be based on an optimal configuration. Or, perhaps, my superficial understanding here leads me to miss some important distinctions.
Llama paper mentioned 135,168 A100 hours for training 13 billion model on 1 trillion tokens, which means ~$150k for lambdalabs on demand instance.
Plus they don't actually have any actually A100s available at the moment (2022-05-17).
CoreWeave is a nice middle ground. You can at least get the A100 machines into a k8s cluster.
this is so US centric :-(
for billions of people, arguably the majority of the world, that’s incorrect
Why?
Dolly's not that great -- I've hit lots of issues using it to be honest .
MosaicML has a nice commercially usable model here: https://www.mosaicml.com/blog/mpt-7b
I think they're one of the leading ones (bias: they're kinda competitors to my employer Anyscale, but you gotta say something's good when it is).
Red Pajama are leading an effort to build a fully open source model similar to LLaMa. https://www.together.xyz/blog/redpajama
I thought the torrents were super active.
This is the fastest I've rolled my eyes in a long time!
I really would ask you to take a second look at the spirit of your comment and think carefully about how much you really understand about the work being done on top of LLMs and if it justifies this kind of response.
If I am an LLM user maybe that’s relevant but prone to being out of date. I’m not going to use this page as the source of truth on that anyways.
Since the article seems to be targeted at developers who use LLMs to e.g. generate Embeddings for semantic search, the title is about as accurate as saying a software engineer is a “keyboard developer” because they use a keyboard.
This is the first time I heard this term, and when I Google search "LLM developer" in an incognito tab, different device, this article is one of the first results.
Seems like we should first establish what exactly is an LLM developer.
> When I was at Google, there was a document put together by Jeff Dean, the legendary engineer, called Numbers every Engineer should know.
The personal plug and appeal to authority of "When I was a Google" is unnecessary. "Numbers every Engineer should know" is public and literally linked there. It's a weird way to start a engineering blog post and makes it feel like marketing of one's resume. Then again, I guess that's what most of these engineering blog posts are nowadays.
Indeed Jeff Dean is a legend and needing to add the "legendary engineer" qualifier detracts from this point. Let these things speak for themselves.