Full threadxtacy·I wonder if the "Lottery ticket hypothesis" work can be applied to this model to further shrink the number of parameters by 10x, to bring it closer to Google's T5 but with higher accuracy?View on HN