https://arxiv.org/abs/2402.17764
The main addition of the new paper seems to be the implementation of optimized and fused kernels using triton, as seen here:
https://github.com/ridgerchu/matmulfreellm/blob/master/mmfre...
This is quite useful, as this should make training this type of LLMs much more efficient.
So this is a ternary weight LLM using quantization aware training (QAT). The activations are quantized to 8 bits. The matmal is still there, but it is multiplying the 8 bit activations by one bit values.
Quantization aware training with low bit weights seems to lead to reduced overfitting by an intrensic tendency to regularize. However, also the model capacity should be reduced compared to a model with the same number of weights and a higher number of bits per weights. It's quite possible that this only becomes apparent after the models have been trained with a significant number of tokens, as LLMs seem to be quite sparse.
Edit: In addition to the QAT they also changed the model architecture to use a linear transformer to reduce reliance on multiplications in the attention mechanism. Thanks to logicchains for pointing this out.