60 karma · joined December 4, 2008
For bigger models (in range of 8B - 70B) the Q4KM is very good, there are no any degradation compared to full FP16 models.
Which table format is better for LLMs? Do you have some insights there?
Recently I have been busy writing the emulator in Golang:
But I believe that most of the data stored in foundation models are just useless for some particular domain. So it's better to forget something, getting really useful info instead.
For example, I use only 6 cores from 10 on my M1 Pro laptop.
And vocabulary is just an array / vector / list - it depends which programming language you use, each has each own terminology for that data structure.
For example LLaMA vocabulary has 32,000 tokens.
https://github.com/gotzmann/llama.go/blob/8cc54ca81e6bfbce25...
[0] https://www.reddit.com/r/LocalLLaMA/comments/140gcn7/new_tok...
https://github.com/ggerganov/llama.cpp/blob/master/examples/...
There's too many schemes right now with 4_0 and 5_1 really popular between LLM geeks.
https://github.com/saharNooby/rwkv.cpp/issues/12
For LLaMA models - yeah, different story.
And some other models have more crazy numbers with even more crazier outliers within them, like you might have a weight of 12.00 between long array of typical small numbers around 0.00
I've read story about attempt to quantize RWKV model into the 4/5 bits which failed short due to the presence of outlier weights.
The author told somewhere that bigger models had worse perplexity because of this.