A 20 Year Old Chipset Workaround Has Been Hurting Modern AMD Linux Systems
phoronix.com
phoronix.com
Kernel : baseline baseline + C2 disabled baseline + patch
Min (MB/s) : 2215.06 33072.10 (+1393.05%) 33016.10 (+1390.52%)
Max (MB/s) : 32938.80 34399.10 34774.50
Median (MB/s) : 32191.80 33476.60 33805.70
AMean (MB/s) : 22448.55 33649.27 (+49.89%) 33865.43 (+50.85%)
AMean Stddev : 17526.70 680.14 880.72
AMean CoefVar : 78.07% 2.02% 2.60%
The origin of the bug appears to be here as well: /* Parse base-64 output produced only by tar test versions
1.13.6 (1999-08-11) through 1.13.11 (1999-08-23).
Support for this will be withdrawn in future releases. */
<http://git.savannah.gnu.org/cgit/tar.git/tree/src/list.c?id=...>GNU tar might randomly make network connections, so sanitize your archive names, I guess..
It's almost worse that rsh isn't installed on most systems.
pledge(2) is the correct model.
The main difference is that when someone balks at the open source code, they can fix it. When someone finds a weird bottleneck on commercial software, our only option is to curse the name of whoever wrote it and move on.
On the other side, I'm always sad / unsettled to see the comments on Phoronix articles like this, the tribalism, the accusations, etc.
The negative comments I saw in the thread are of the style below.
> That also tells us that AMD has not bothered to fix this for 20 years. On the other hand, I welcome their stepped-up efforts to improve the Linux experience for their end-users. There are still some areas where more participating could yield some benefits though, e.g. glibc or better upstream compiler support/tuning.
and
> It seems Linux competition is so slow AMD didn't even notice such suboptimal performance.
Unsettled? This is milquetoast criticism at best.
Link to the forum below in case I missed something. I've never visited the forum before and am not deeply involved in Linux, other than happily using it for my home computer.
https://www.phoronix.com/forums/forum/phoronix/latest-phoron...
On the face of it, I'd expect the major players (RedHat/MAGMA) to be profiling the heck out of common SW/HW combos, but I can imagine there are a lot of scenarios they don't cover, and I'm not sure who else would.
So, could the impact only be substantial in very specific workloads? Because otherwise someone would have found it already?
My current suspicion is it's somehow related to power saving, because I can't reproduce it after a fresh reboot. It only starts happening once I go through a few power states (e.g. by closing the laptop lid).
Since you mentioned a mobo I assume you're on a desktop. Do you use sleep/hibernate?
Although sounds like your problem could also be caused by I/O saturation?
The issue is a dummy op used to work around an issue in old hardware gets counted as C-State residency, and "A large C-State residency value can prime the cpuidle governor to recommend a deeper C-State". So this dummy op used during switching C States makes the CPU look more idle than it is, causing more low power mode than there should be.
So umm e.g. long compilation jobs of complex software? What else?