The paper talks about parameter partitioning and overlapped communication, but doesn't actually give many details on how those things happen.
The library appears to be an implementation of some common algos for solving the 'pebble game,' as explained decently here: https://medium.com/tensorflow/fitting-larger-networks-into-m...
The essential point is that:
(1) model parallelism is hard to do and has historically been done manually to scale wide models across GPUs
(2) inter-GPU I/O is expensive for vanilla data-parallel jobs (that typically use naive mirroring strategies)
(3) researchers have figured out now how to 'compile' a deep model so that layers span GPUs and save on both memory usage and I/O
(4) so scaling wide models is still hard, but now we have better tools for deep models
Existing all-reduce-based data-parallel problems have already been well-studied (see e.g. https://people.eecs.berkeley.edu/~jfc/papers/14/Kylix.pdf ), so it's really nice to see gains through new techniques.
Definitely like seeing this 'compilation' being wrapped up into a library. Just wish they did a better job of communicating key ideas.