Task Failed Successfully: Saturating NIC and Disk Bandwidth
blog.mrcroxx.com
blog.mrcroxx.com
After I identified the TLB misses and confirmed that huge pages were effective, I noticed that there were still many suspicious points in the flame graph. Interestingly, my agent attributed the effectiveness of huge pages to those suspicious points, which turned out to be unrelated to the bottleneck at the time. That sparked my curiosity.
The structure of this blog post was mainly chosen to make the story easier to follow, while also covering the various issues I investigated in depth along the way.
In fact, I recently switched from my previous job building data infrastructure on cloud to an HPC-related role, so I am still not very familiar with some of the mature practices and established conclusions in the HPC world.
So thank you very much for your suggestions. I also hope to learn about more and better methods that can help people identify root causes more quickly and accurately in complex scenarios.
Just throwing out some ideas, obviously the best solution is the one that you already have working :)
The short version: READ_FIXED fixed the obvious per-I/O GUP overhead in a small demo, but the larger deployment still got stuck at roughly half of line rate. After ruling out io-wq backlog, request splitting, fd lookup, and CRC arithmetic, the actual wall turned out to be dTLB misses from scanning 1,028 KiB buffers backed by 4 KiB pages. Moving the read arena to hugepages brought the system close to NIC saturation.
The funny part is that an AI agent suggested hugepages early and got the optimization right, but its explanation was wrong. This post is mostly about reconstructing the evidence for why it worked.
I’d be very interested in feedback from people who have used AI to debug performance issues in a complex system.
So anyone familiar with the space could have suggested something like that without knowing the details of the problem. Hence it is not useful advice IMO.
That aside, the blog post was really cool to read and a instant favorite, wish there were more english posts on the blog.
Especially like the hardware limit based expectations, detailed measurements and the writing style.
Finally, thank you very much for your appreciation, which means a lot to me. Previously, I was working on open-source projects, but now that I’ve changed jobs, I may not have the same amount of energy to contribute to open-source code as before. However, I think blogging might be a new way for me to contribute. I hope I can keep it up.
(My English writing skills are poor, so I wrote in Chinese and used AI to translate it; I hope you don’t mind.)