2. Otherwise, most of these are obvious optimizations. One widely popular optimization (that I consider non-obvious) is ZeRO-Offload, in particularly the gradient sharding scheme (although once learned, the sharding itself is pretty straightforward, just a bit chatty). One thing I think undervalued these years though is Alex's "One weird trick": https://arxiv.org/abs/1404.5997. This scheme is much convoluted but very effective when training large MLP. It is not popular probably because a). the implementation is non-obvious; b). large MLP fall out of fashion quickly, and the computation shape for transformer looks probably very different from the MLP its originally trying to solve (with 4096 activations per layer).