Just as a quick example (there are lots of other examples), I currently am working on a model that does significant tensor-level parallelism. The input to the model is huge, so we replicate the weights for the encoder part of the model over 12 A100s, and then run an all-reduce over the outputs before then distributing that reduced tensor to another set of GPUs for the next set of operations. In this setup we need significant bandwidth not only between every device in a node but also across nodes (infiniband).