"Mixture of Experts (Yiet al., 2023) have a sparse structure in their feed forward layer. This property can be used to combine with our method for enabling larger MoEs on device."
Assuming this implementation would allow for running Mixtral 8x7b on a 16Gb M1, I'm happy.