Not actually, you might see here how bigger models have much worse perplexity with 4bit due to the weight outliers:
https://github.com/saharNooby/rwkv.cpp/issues/12
For LLaMA models - yeah, different story.
https://github.com/saharNooby/rwkv.cpp/issues/12
For LLaMA models - yeah, different story.
No comments yet.