Interesting, how do you use -0 in the add, then? Is -0+1-1 a 0 or a -0?
> Could the additional -0 carry some pseudo-gradient information
It looks like training was done on fp32 or bf16. Low-bit quantization is approximated with STE during training. I'd expect training itself cause each point to "polarize" towards 1 or -1.
> 2-bit quantizations being proposed
Symmetric (i.e. without 0) exponential values were pretty popular IIRC.