The 4 bit quantization performs well, though. Does your 1 bit version?
However I can see fractional bits (via binary representations) and larger models happening first before that compression step.
And then we have the sub-bit range..... ;DDDD
So we have numbers on PTB original perplexity 8.79 quantized 9.68, already 10% worse. And PPL reported per token I suppose? Because word PPL for PTB must be around 20, not less than 10.
Any numbers on more complex tasks then? like QA?