As someone who only has a passing understanding of parallel training, I always found it wonderful that averaging gradients works at all. It seems non-intuitive to me.
Pretty cool, imho, that there are finding better ways to train in parallel.
Pretty cool, imho, that there are finding better ways to train in parallel.
This is indeed super cool !