Why async gradient update doesn't get popular in LLM community?
github.com
github.com
However, both the Megatron-LM[3] and the DeepSpeed[4] don't use pipedream-2bw scheduling. Could anyone share me some insights or ideas about why such an efficient scheduling scheme doesn't get popular in the LLM pretraining community? Does it suffer convergence/accuracy issue in practice? Or are there any other concerns that blocking it become the default / most popular pipeline parallelism scheduling?
[1]: https://arxiv.org/abs/2006.09503
[2]: https://arxiv.org/abs/2101.06840