Horovod: Distributed Training Framework for TensorFlow, Keras, and PyTorch
github.com
github.com
PyTorch? You mean Theano, right?
They use a distributed training model that utilizes parameter servers, which scales nowhere near Horovod's mpi solution.
Even for single-machine-multi-gpu solutions, only now in Tensorflow 1.8 is pure tensorflow as fast as Horovod with it's estimator MirroredStrategy. If you watch Tensorflow dev days 2018, the devs say they're working on bringing something like Horovod to pure Tensorflow
in my limited experience with horovod, horovod is most useful when youre running large clusters of workers/ps. in those situations, you typically have to manually find the appropriate balance of workers/ps. (otherwise youd run into blocking or network saturation issues.) horovod addresses this issue with their ring allreduce implementation.
having said all of that, im sticking with distributed tf for now.
Sure it's nice to "achieve 90% scaling efficiency" in images/sec, but images/sec alone doesn't get you where you want to go. Increased accuracy per unit wallclock time is what you want.
yah, especially for a framework for distributed learning!