The performance gap is shown to be due to hardware offloading, not due to congestion control, in the arxiv paper above.
And it’s often incorrect on x64 PCs when IOMMU access is appropriately segmented. See also e.g. Thunderclap: https://www.ndss-symposium.org/wp-content/uploads/ndss2019_0...
It may still be true in some cases, but it shouldn’t be taken for granted that it’s always true.
Although I think for Intel CPUs the mmunuded to be disabled for years because their iGPU driver could not work with it. I hope things have improved with the Xe GPUs.