So I guess my question is - what use case is there for a huge truck that goes 200mph and take 4 trips, when you could just buy 16 regular trucks, and move your apartment in the same amount of time at half the cost.
So I guess my question is - what use case is there for a huge truck that goes 200mph and take 4 trips, when you could just buy 16 regular trucks, and move your apartment in the same amount of time at half the cost.
'The street finds it's own uses for things' is the well known Gibson adage, I and typically it's a comment aimed low. But our entire era of amazing computing began with the Gang of Nine enabling lowness in a degree such that it quickly became the highest tech, the best. Sure you can still buy a mainframe & they have amazing feats but it's not where the value is, but and the value is where it is because possibility was unchained, I unleashed from corporate dominion, and spread wide. I think we can find amazing new futures with CXL & mad bandwidth connectivity.
The other approach, which you do when models themselves are massive, is model parallelism. You split it into multiple parts that run on different nodes.
In both cases, you need to distribute weight updates through the network although the traffic patterns can be different.
To maximize the performance in both scenarios, systems designers optimize for all-reduce and bisection bandwidth.
There are also other tricks, for example the TPUv4 ICI network is optically switched, and it is configured when a workload starts to maximize bandwidth for the requested topology ("twisting the torus" in the published paper).