There's also research showing that the perplexity reduction is less at higher parameter counts. E.g. a 65b parameter model barely has any impact at all when reducing from 16bit to 4bit
https://github.com/saharNooby/rwkv.cpp/issues/12
For LLaMA models - yeah, different story.