Personally, I think that the RLHF does make a big difference but maybe it's a bug in the quantization code as suggested up thread.
Personally, I think that the RLHF does make a big difference but maybe it's a bug in the quantization code as suggested up thread.
It seems like if somebody figured out the “correct” way to quantize the 7b weights it would make way more sense to just torrent the output rather than distribute a fixed program.
i) Distributing large files through torrents is slightly annoying if you don't already happen to have a seedbox
ii) People are still messing around with quantization settings, they might think that they are a few days away from a much better version
iii) No one wants to be sued by Meta. I think the risk is pretty small but not zero.
https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/READ... says llama-13B takes 42GB and 33B takes more than 64GB...