> We proceeded without DCQCN for our 400G deployments. At this time, we have had over a year of experience with just PFC for flow control, without any other transport-level congestion control. We have observed stable performance and lack of persistent congestion for training collectives.
https://engineering.fb.com/2024/08/05/data-center-engineerin...
But now I am curious with the distribution of observed window sizes is in the wild.
Edit: I'd bet the simpler protocol is more vulnerable to various spoofing attacks though.
Edit2: Lol I hope the frame IDs are for illustrative purposes only - https://chipsandcheese.com/2024/08/27/teslas-ttpoe-at-hot-ch...
Such ideas are, however, worth revisiting when the workload is unique enough (in this case, it is), and the performance gains are so big enough...
Multiple parties communicate at the same time? Lower number priority electrically could pull the voltage low, dominating the transmission.
That way, priority messages always get through with no overhead or central communication required.
The technical issue is that you would need global arbitration to ensure that the _goodput_ (useful bytes delivered) is optimal. With training across 32k GPUs and more these days, global arbitration to ensure the correct packets are prioritised is going to be very difficult. If you are sending more traffic than the receiver's link capacity, packets _will_ get dropped, and it's suboptimal to transmit those dropped packets into the network as they waste link capacity elsewhere (upstream) within the network.
This is a protocol between compute nodes in a data center, it's layer 2 so there is no way to reach this over the internet.
But, point taken.