Use `nproc` and not grep /proc/cpuinfo
flamingspork.com
flamingspork.com
np = sysconf(_SC_NPROCESSORS_ONLN);
However I'm interested in whether this is cgroups- (hence container-) aware or not. I'm not actually sure, and it will affect some of my software when running in a container.It would be ideal to have a standard utility or library that would also check that.
If you want to understand more about these problems this article has some good background, although they fixed a real bug in the scheduler here the background will give you the general idea and it's still a problem to some extent if all of those threads/processes are trying to contend on a single shared lock.. which is precisely what happens with multi threaded Python programs (they sometimes contend on the GIL) that use data science modules which do a lot of processing in C without the GIL lock held and only take it occasionally. However the problem comes if you have 12 threads consuming your 4 threads worth of CPU time they'll end up block for 5-10 or more milliseconds at a time where they may well be holding that lock and preventing the other threads from running.
You may then ask why not reduce the CPU set for those containers, the problem is there is no way (at least currently) to say 'this process can only run on 4 CPUs, but any of them'. You have to reduce the CPU set down to 4 specific CPUs and now if all containers scheduled to those specific 4 CPUs are busy you will be losing time to those other containers when a different 4 CPUs may be totally idle. It would be nice to advance the scheduler to support such a concept without using the quotas specifically.
As a final side note.. LXD cheats here. It uses 'lxcfs' to bind mount onto /proc/cpuinfo and modifies the output to only show the number of CPU cores you actually have. Which suffices for most programs.
https://kubernetes.io/docs/tasks/administer-cluster/cpu-mana...
A somewhat annoying thing about this option is that even though it can create cpusets for every container individually, it only gets enabled in case all containers in the pod use an integer number of CPUs, which seems unnecessary.
I think this is a jumping to conclusions a bit. You’d need to be reading /proc/cpuinfo in a fairly tight loop in order for the time spend doing any context switching to dominate the runtime cost of a program. Is this a common practice in some types of computation? Or am I misunderstanding this quote from the article?
> In this case, if you base your number of threads off grepping lscpu you take another dependency (on the util-linux package), which isn’t needed. You also get the wrong answer, as you do by grepping /proc/cpuinfo. So, what this will end up doing is just increase the number of context switches, possibly also adding a performance degradation.
The idea is that a common way to size an application that exploits concurrency is to base the number of application threads on the hardware threads available. This portion of the article is saying that if you do so on the basis of `lscpu` or grepping `/proc/cpuinfo` you could mis-size relative to the actual number of cores available to schedule the threads of your application on. Doing so will lead to excessive context switching for your application, and that can hurt.
There may be 10 CPUs in the machine, but if your in your container which is restricted to using only 1 of them, it's no use launching 10 threads in that container over just having one.
getconf _NPROCESSORS_ONLN
% taskset 1 nproc
1
% taskset 1 getconf _NPROCESSORS_ONLN
8> expert MAKEFLAGS=-j$(nproc)
make -j without a number should do that, instead of spawning processes like there’s no tomorrow.
I had a makefile with a race condition that triggered only sporadically. However, giving it more threads significantly increased its rate of occurrence.
make -j$(($(nproc)+2))Or prefer ninja overall.
OMP_NUM_THREADS and OMP_THREAD_LIMIT are in principle extra limitations, not the main source of information, right ?
$ strace nproc
sched_getaffinity(0, 128, [0, 1, 2, ..., 78, 79]) = 40
fstat(1, {st_mode=S_IFCHR|0600, st_rdev=makedev(0x88, 0xf), ...}) = 0
write(1, "80\n", 324) = 3
exit_group(0) = ?
+++ exited with 0 +++with a big.LITTLE system it'll really depend what you're doing. Just compiling something? It's probably fine (if all cores can be active at once). If you're wanting more consistent performance for some calculations? You might need to think a bit more.
$ nproc
ksh: nproc: not found
$ uname -a
OpenBSD obsd64.example.xyz 6.8 GENERIC.MP#828 arm64
$ sysctl hw.ncpu
hw.ncpu=2