High System Load with Low CPU Utilization on Linux? (2020)
tanelpoder.com
tanelpoder.com
Regarding the exact case, there is a slightly deeper issue. XFS enqueues inode changes to the journal buffers twice: the mtime change is scheduled prior to the actual data being written, and the inode with the updated file size is placed in the journal buffers just after. If the drive is overloaded, the relatively tiny (just a few megs) journal buffers may overflow with mtime changes, and the file system becomes pathologically synchronous. However, since 4.1something, XFS supports the `lazytime` mounting option that delays the mtime updates until a more substantial change is written. Without it, the journal queue fills up at roughly the speed of your write() calls; with it, at the pace of the actual data hitting the disk, so even in highly congested conditions your application can write asynchronously -- that is, until dirty_ratio stops your system dead in its tracks.
When I first tried this, I was prepared to hard-boot since I was almost sure it would make my desktop unusable, but it didn't. I can even play some fairly cpu- and gpu-intensive games without too many hiccups while this is going on. If I wasn't paying attention, I probably wouldn't know they were running.
Great article but this summary was zero surprise. I've only ever seen high load from disk I/O. When I was first clicking on the link I thought to myself "well, it's disk I/O, but let's see how we get to the punchline"
The other reason is that I've troubleshooted plenty of Linux load spike problems that are about CPU demand spikes only, usually due to some spinlock that gets held unusually long or some interrupt storm issue or some sort of a "database logon storm" due to connection pools in the app server suddenly creating thousands of additional DB connections...
I'm liking this project https://github.com/facebookincubator/below
It's packaged in Fedora.
Edit: Nevermind. I skipped over the key paragraph here:
> By default, pSnapper replaces any digits in the task’s comm field before aggregating (the comm2 field would leave them intact). Now it’s easy to see that our extreme system load spike was caused by a large number of kworker kernel threads (with “root” as process owner). So this is not about some userland daemon running under root, but a kernel problem.
Edit: I understand that PSI was created as an easy (and cheap) to query metric to see if your Android mobile device has a problem (and some apps need to be killed) and not really for full blown troubleshooting drilldown of server workloads.