Understanding Linux CPU Load - when should you be worried?
blog.scoutapp.com
blog.scoutapp.com
This article entirely ignores the number of processes blocked on I/O. A load average exceeding the number of CPUs (cores, whatever) does not automatically mean the CPUs are overloaded.
How did this end up on HN? You have equally good articles on load at Wikipedia.
1. I can't comment on the original article. Are comments closed, or am I dumb?
2. The author seems to have assumed a web server responding to bursty traffic. Several people have pointed out workloads to which the 0.7 heuristic doesn't apply - compute servers, I/O bound servers, compile jobs, desktops. He should have stated that assumption up front.
3. Hyperthreads. For purposes of load monitoring, you should be counting the number of threads, not the number of cores. Yes, hyperthreads are slower than cores, but that doesn't matter. The load average is the ratio of work available to work being done (oversimplified, I know), and, as such, it's scaled to the actual throughput of the threads available.
Fortunately, the author suggested counting CPUs by reading /proc/cpuinfo, and /proc/cpuinfo lists threads, not cores. So those two errors cancel out. (-:
It was posted in 2009, so the comments are probably closed :-)
2) They make that assumption because they are selling web server monitoring tools, if you don't have a web server you aren't their target user :-)
3) Now you actually want to talk about real details about monitoring performance and the goal of this particular article is to sell their product to people who run web servers and probably don't want to delve too deeply into actual performance analysis.
Hope that helps.
So it's not that SMT pipelines are slower, it's just that they share resources with the other SMT pipelines.
[1] Simultaneous multithreading (SMT), http://en.wikipedia.org/wiki/Simultaneous_multithreading, is the generic name for what Intel calls hyperthreading.
If (# CPU Cores / Load) > 1, shit has hit the fan
I disagree with 0.7 being the starting point for investigation on extraneous load, but you should be more worried about changes in 1st or 2nd moments in the load (velocity and acceleration), which as analogy on the link, you don't care too much about steady traffic, it's when traffic starts bursting at the seams.
Having a machine running at 0.75 load for a shared machined (say a development database) might actually mean your resources are actually being consumed regularly. Albeit seeing that average load climb slowly towards ~1.0 means you need to fix it before the pipes clog shut.
Although, it is much easier to go from 1.0 to 2.0, as opposed to 0 to 1.0 on such a cpu, because the cpu can't handle much more work before it gets overloaded.
For example, in the classic apache dos failure state, you wind up with enough apache processes that some of them are forced out of memory, and then you start swapping and fall over. What you'll see then is a really high load average and long page load times. Looking at something like vmstat 1 or top, or iotop you can see if it looks like memory, disk, or something else.
From your description, you probably have enough cpu resources to saturate something else. Maybe the DB server, maybe your memory. When that happens, your processes stack up and your load average rises. It looks like you don't have neough CPU, but that's probably not it.
This is because traffic is rarely "smooth." Even if you say the link is operating at 90% utilisation, you usually refer to the average. Pushing a system to high load can lead to instabilities and unpredictable performance.
Caring about 0.7 load means you've got some capacity left if traffic does burst. You generally don't have warning, so having a healthy amount of CPU left available is generally good.
system_profiler SPHardwareDataType
Hardware:
Hardware Overview:
Model Name: MacBook Pro
Model Identifier: MacBookPro5,1
Processor Name: Intel Core 2 Duo
Processor Speed: 2.4 GHz
Number of Processors: 1
Total Number of Cores: 2
L2 Cache: 3 MB
Memory: 4 GB
Bus Speed: 1.07 GHz
Boot ROM Version: MBP51.007E.B05
SMC Version (system): 1.41f2
Serial Number (system): [snip]
Hardware UUID: [snip]
Sudden Motion Sensor:
State: EnabledPut an "ionice -c 3" on that job and you probably won't notice the performance effect on Firefox anymore. (You probably don't need conventional "nice" because compile jobs tend to get their priorities dropped anyhow because they are using a lot of CPU without yielding, but dropping the scheduler that hint can still be helpful in some cases.)
(Annoyingly, unlike nice, ionice requires the specification of the class; I wish it would just default to -c 3 like nice has a reasonable default.)
schedtool -B -e ionice -c3 make -j10
That's the idiom I commonly use on long large compile jobs. means that anything else will always get the cpu or io time (maybe with some increased latency, which usually isn't too bad) that it needs while all idle time is taken up by the larger compile job. This makes for a very happy desktop system when upgrading (gentoo user here).https://lwn.net/Articles/418739/
Seriously. Nobody _ever_ does "nice make", unless they are seriously repressed beta-males (eg MIS people who get shouted at when they do system maintenance unless they hide in dark corners and don't get discovered). It just doesn't happen.
:)
Now, "make -j" on Linux source tree is not something most machines recover from.
ftp://crisp.dyndns-server.com/pub/release/website/dtrace/