While I think TF probably has a better distributed story, it's still not great (several months of experience getting this to work), and very very few people can justify distributing computation across 10K CPUs/GPUs - in my experience, it turns out most async SGD algorithms don't really work very well.
One place I see this actually working is using SVI/EM and storing local parameters across the network - something maybe 1 GPU can't handle for extremely large models.
I’m sorry to hear that you’ve not had a pleasant experience with tf.data. One of the doc-related criticisms we’ve heard is that they aim for broad coverage, rather than being examples you can drop in to your project and run straight away. We’re trying to address that with more tutorials and blog posts, and it’d be great if you started a blog on that topic to help out the community!
If there are other areas where we could improve, I’d be delighted to hear suggestions (and accept PRs).
More than documentation, I would argue that TF especially tf.data lacks a tracing tool that would let a user quickly debug how data is being transformed and if there are any obvious ways to speed up. E.g. image_load -> cast -> resize vs image_load -> resize -> cast had different behavior and lead to hard to identify bugs. For tf.data prefetch which ends up being key to improving speed yet its is not documented, the only way I actually found out about it was by reading your TF.Data presentation.
This may not seem useful this conventional training, where you usually work with a fixed amount of samples you know beforehand. But there may be cases where this is not true (for instance, in some special cases of augmentation) - the streaming part is useful but then you must use this caching trick.
But I agree API naming is not stellar, or at least should come with better documentation.