Can't find documentation online yet for this year's H100 systems, but here's a schematic for an A100 server: https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-...
There 2 GPUs, 2 NICs and 1 NVME are connected via one PCIe switch.
The torus topology makes sense for ring algorithms: Allreduce, Allgather, Reducescatter. For purely data parallel training you could put all model replicas into the same ring (although Nvidia also uses hierarchical algorithms that benefit from lower lately). With added model parallelism one will need smaller rings running concurrently. I guess the TPU cluster layout will then put constraints on the most efficient model architectures (as does the network topology of a GPU cluster).