Where Linux's load average comes from in the kernel
utcc.utoronto.ca
utcc.utoronto.ca
For those that are interested:
https://tanelpoder.com/posts/high-system-load-low-cpu-utiliz...
But yeah, even that doesn't clearly tell you what amount of active threads is normal for your workload/configuration. Or you can just save the thread count info (grouped by thread state) and save/graph it over time, to have a better idea of what's normal. The difference of this approach vs just graphing Linux-reported load is that you'll get a breakdown of how much of it was CPU demand vs sync I/O demand. Or you could even break the the thread states down further, by syscall and WCHAN, for example. This will give you a much more detailed idea of why the load is high, not just that it is high.
(Like I mentioned in the article), I even have a tool for sampling /proc at regular intervals (1 Hz by default) and saving it into hourly CSV files, so you can easily go "back in time" and zoom into short spikes and see what was going on:
$ xcapture
0xTools xcapture v1.0 by Tanel Poder [https://0x.tools]
Sampling /proc...
DATE TIME PID TID USERNAME ST COMMAND SYSCALL WCHAN
2020-10-17 12:01:50.583 6404 7524 mysql R (mysqld) fsync wait_on_page_bit
2020-10-17 12:01:50.583 6404 8944 mysql D (mysqld) fsync wait_on_page_bit
2020-10-17 12:01:50.583 6404 8946 mysql D (mysqld) fsync wait_on_page_bit
2020-10-17 12:01:50.583 6404 76046 mysql D (mysqld) fsync wait_on_page_bit
2020-10-17 12:01:50.583 6404 76811 mysql D (mysqld) fdatasync xfs_log_force_lsn
2020-10-17 12:01:50.583 6404 76815 mysql D (mysqld) fsync blkdev_issue_flush
2020-10-17 12:01:50.583 8803 8803 root R (md10_resync) [running] 0
DATE TIME PID TID USERNAME ST COMMAND SYSCALL WCHAN
2020-10-17 12:01:51.623 6404 7521 mysql D (mysqld) pwrite64 xfs_file_buffered_aio_write
2020-10-17 12:01:51.623 6404 7524 mysql D (mysqld) fsync xfs_log_force_lsn
.../*
* kernel/sched/loadavg.c
* This file contains the magic bits required to compute the global loadavg
* figure. Its a silly number but people think its important. We go through
* great pains to make it work on big machines and tickless kernels.
*/
And on the topside, if you block on the io_submit and friends, async I/O counts towards the PSI and load average again. Ie, when you can't do useful work because the device queue is full, for example.
So it's not as much of a shortcoming as you think, because if the system is I/O exhausted, async I/O processes will quickly contribute to load as well once the I/O queues are nearly full or full.
So, the I/O component of Linux system load and PSI do not include the async I/O waiting threads that decide to synchronoysly wait for asynchronously submitted I/O requests :-)
I see MySQL is using libaio in some places too (the io_submit/io_getevents show up in syscalls), but I haven't looked deeper into whether the getevents are "willing to wait" or not. But any application using libaio, that at some point uses io_getevents() in "willing to wait" mode, would be affected by this discrepancy.
The crucial point there is that while the application is slow, the system remains as reactive as the PSI number indicates. Just the application didn't.
But I'm talking about the cases where the app is done with the other work and needs to ensure that the previous write is persisted (for example to a write ahead log), before moving on? It will have to deliberately wait for I/O completion now, thus will run io_getevents() with the "willing to wait" mode enabled. I see this in Oracle database world all the time and MySQL world occasionally too:
- Not seeing a high number of threads in D state or PSI IO figure does not mean that there are no threads waiting for I/O
- Seeing a high number of threads in D state and high PSI IO figure means that there are threads waiting for I/O (but possibly even more, if you'd look into the ones in S state, but in io_getevents syscall, with WCHAN=read_events.
Then your app, like Oracle, simply issues a fence and can be done with it. If the app crashes before the fence is persisted, then it must be able to resume work from a previous one (a simple example would be that between WAL checkpoints, a fence is issued). The app won't have to worry that only some specific writes completed vs some others not beyond what the fence permitted. Additionally a good mechanism might be a call to wait for a fence to be persisted to disk.
Simple example for the WAL use case:
1. Write new transaction to WAL
2. Issue an IO fence for the WAL file
3. Write new data to database file
4. Wait for Fence from 2
5. Return success
This is roughly equivalent to the synchronous example: 1. Write new transaction to WAL file
2. Issue fsync
3. Write new data to database file
4. Return Success
In case writes to the database file are lost, you can recover from the WAL (as intended). Notably Step 4 of the Async Example is a case where the thread is waiting but it can do other useful work while that is happening. The same thread can offload the work and simply issue more IO in the meanwhile and return success to the client once it sees the correct wait-for-fence returning in the async queue. And it won't have to wait for AIO event completion/checkpoints like currently, that make the system load non-indicative of app load (though frankly, system load is never indicative of app load, apache2 doesn't increase load if it runs out of workers).1. db file async I/O submit
2. db file parallel write
(the 2nd one has such a name for historical reasons, but it's the I/O reaping/completion check timing, not the submission).
Of course, 100% CPU use also isn't a "problem", it's just a fact. PSI is the thing that actually reveals problems. I'm not aware of a Windows equivalent.
Low-level hardware block usage generally isn't accounted for that way, but CPUs can keep track of performance counter events tracking how many of each operation are performed and how many cycles each subunit is active. The OS can periodically read that data, but it's only used for low level metrics about how efficient the code is rather than for calculating overall CPU utilization.
There are separate measurements about GPU and hard disk utilization, which are either determined by when driver queues work on the device and receives a notification that the device is done, or by the device keeping track of what work it's switching between and reporting it periodically.
So it may be the case that you showed people the door, simply because their experience was based on non-Linux operating systems.
When I've been on the employer side of the interview table my priorities are: does this person have a rudimentary knowledge of the required subjects and what is their problem solving process like. Basically do they know enough to handle the majority of day to day things and know where to look and who to ask when something is beyond their knowledge. Arcane trivia is of little-to-no interest to me because nobody is going to memorize every single detail. There will always be surprises. And, yes, implementation details of the load average is arcane trivia – that's precisely why this article is/was on the front page of HN.
With the scenario you've put forth it's absolutely possible to hit the ground running without having to know the gory details of how the Linux kernel calculates load average. Surely you're already monitoring I/O activity and CPU usage alongside load average. And if you're not (tsk tsk) a competent candidate would know where to look for current system information, and how to respond to the information (e.g. is this elevated load average worth being concerned about in the first place?).
> Best way to show you are a motivated, quick study is by being a motivated, quick study.
The best advice I can give is to look for successes and not failures in your candidates. Asking something tantamount to a trick question. Asking them to work through a scenario where the load average is high and the CPU utilization is low is looking for success.
It's hard to get generalized knowledge being a quick study, it's a lot easier to get specific knowledge to work on an actual issue being a motivated, quick study. So if that's really what you're looking for, you've got to be OK with using a search engine and manual pages during the interview.
Asynchronous I/O completion checks (libaio's io_getevents) that are willing to wait for I/O completion, will not contribute to Linux system load (threads in S state).
Asynchronous I/O submissions (libaio's io_submit) either quickly submit their I/O (a small amount of time in R state) OR get stuck in io_submit() if the underlying block device I/O queue is full. When io_submit() gets stuck, then you're sleeping in D mode, thus contributing to system load again.
https://utcc.utoronto.ca/~cks/space/blog/unix/ManyLoadAverag...
as one of many possible theoretical scenarios.
I can't count the number of times this has helped solve a mysterious behavior. Atop is king.
"Suppose, not hypothetically, that you have a machine that periodically has its load average briefly soar to relatively absurd levels for no obvious reason; the machine is normally at, say, 0.5 load average but briefly spikes to 10 or 15 a number of times a day. "
$ sudo perf sched record -- sleep 5 && sudo perf sched latency
seems to do the trick. I'm not even a perf engineer just did a simple google search. Though, that makes for pretty crummy blog content.
I think the author didn't realize that the unit of load is just "number of threads" doing _something_ (wanting to be on CPU, running on CPU or waiting for I/O in D state on Linux, just CPU stuff on other Unixes).
Load average is just "average number of threads" doing something over last 1,5,15 minutes.
So if just the single-number averaged over multiple minutes is not good enough for drilling down into your load spikes, then you just go look into the data source (not necessarily source code) yourself. Just use ps or /proc filesystem to list the _number of threads_ that are currently either in R or D state. That's your system load at the current moment. If you want some summary/average over 10 seconds, run the same commant 100x in a row (and sleep a bit in the between) and count all threads in R & D state (and then divide by 100 to normalize it to an average).
It's basically sampling-based profiling of Linux thread states.