Why do you need all-to-all? With deep networks you could just put different layers on different GPUs, then run the forward and backward passes asynchronously. They would only need to communicate in a chain-like manner.
Just as a quick example (there are lots of other examples), I currently am working on a model that does significant tensor-level parallelism. The input to the model is huge, so we replicate the weights for the encoder part of the model over 12 A100s, and then run an all-reduce over the outputs before then distributing that reduced tensor to another set of GPUs for the next set of operations. In this setup we need significant bandwidth not only between every device in a node but also across nodes (infiniband).
Isn’t the simpler example data parallel training? That’s a reduce and broadcast.
You’re correct, I was going to add on to my answer that this is combined with DDP for a “3D” style parallelism that could specifically benefit from the TPU’s network topology, but by the time I got done fixing all the typos (and still missing a few) from writing it on my phone I completely forgot :)
To be honest, I think regulators should have stepped in to stop the acquisition of Melannox (Infiniband) by NVidia. There are basically no non-NVidia options for Infiniband now. (Infiniband being about half an order of magnitude ahead of Ethernet at any one time, with lower overhead and lower cost per unit bandwidth. Often the basis for high bandwidth Ethernet cards as well. And ostensibly an open standard.)
I agree that I’m uncomfortable with the consolidation of so much of the HPC stack under Nvidia (and would be 10x more uncomfortable if the ARM acquisition went through), but I don’t think there was really a monopoly there. I think the real “market” is networking, not just infiniband, and in that sense there is a ton of competition, just no competition that hits the performance characteristics that Nvidia/Melanox achieve.