Forcing the CPU affinity can make a monothreaded process run 2-3x faster
klaig.blogspot.com
klaig.blogspot.com
In my experience writing high performance VFX software, the Linux kernel's scheduler has been the best of all major OSs in terms of balancing threads since around 2.6.35.
OS X is the worst, it bounces threads all over the place, and on top of that, thread_policy_set() on OS X is only a hint, so often OS X will ignore affinity settings anyway.
I've seen this for example encoding AAC audio. I have an 8-core system but the encoder is single-threaded but the Windows scheduler still spreads the process out over 8 CPUs. Wouldn't it be better to stick on 1 core causing cache hits to be higher?
In the past the power management on Linux was done by a user-space program, I'm not sure why that was changed.
But scheduling "hinting" sure, like saying "ok, keep this thread to one processor only"
Still, some of these configs may cause some kind of 'denial of service' on the system if misused (like the 'nice' command), so they're usually limited.
But the main problem is that workloads are very complex, and often changing. A single userspace program is not even aware of other programs running on the machine, nor how the user thinks they should be prioritised.
If you are compiling something, do you want to to finish ASAP or a bit slower in the background without desktop sluggishness? What priority should a minimised browser window have at the same time? What if that browser window is also playing music from youtube?
At the moment, the scheduling hints that users can figure out is marking some processes as low-priority background tasks. Optimal scheduler tuning is too complex, because starting or quitting any application can completely change the optimal resource usage, and users only have a vague idea of what they expect from the scheduler ("everything should be fast").
They do try, at least the one on Linux does. But the OS can't know what your intentions are when you fire up a thread.
http://www.kernel.org/doc/Documentation/scheduler/sched-desi...
I don't know how this deals with distributing tasks across multiple processors however; the basic idea was to improve on the fairly naive runqueue implementation in 2.4 and prior.
- Your AAC encoder runs UN CPU #1.
- Your AAC encoder stalls on I/O.
- The scheduler picks some other thread to run on CPU #1.
- The I/O request completes.
- A third thread, running on CPU #2, blocks.
- Your AAC encoder is the first waiting thread.
Should the scheduler: - run your AAC encoder on CPU #2?
- run your AAC encoder on CPU #1,
and move the program happily running there to #2?
- wait until CPU #1 becomes available before running your AAC encoder again?
Keep in mind that 'the program happily running there' could very well be my AAC encoder.Variations include the case where, by the time your I/O completes, CPU's #1 and #3 are available, but #1 is asleep. Should the scheduler wake it, so that your thread can stay on the CPU?
Your AAC encoder may be the most important process alive for you, but how is the scheduler to know that?
(Running processes you deem important at a nice level might help, but I do not know enough about current schedulers to know about that)
This is rare on Windows client machines. Does anyone know if desktop Windows' scheduler uses different rules to their server OS to exploit this?
Server OSes typically use different schedulers or scheduler settings than desktop ones; Windows is not different. This starts with using larger time quantums. http://download.microsoft.com/download/1/4/0/14045A9E-C978-4...:
"On client versions of Windows, threads run by default for 2 clock intervals; on server systems, by default, a thread runs for 12 clock intervals"
The windows scheduler also is aware of the GUI, and raises priority of threads handling the user interface. Some things the kernel and/or user mode code do:
"Threads that own windows receive an additional boost of 2 when they wake up because of windowing activity such as the arrival of window messages . The windowing system (Win32k .sys) applies this boost when it calls KeSetEvent to set an event used to wake up a GUI thread ."
"Client versions of Windows also include another pseudo-boosting mechanism that occurs during multimedia playback . Unlike the other priority boosts, which are applied directly by kernel code, multimedia playback boosts are actually managed by a user-mode service called the MultiMedia Class Scheduler Service (MMCSS), but they are not really boosts—the service merely sets new base priorities for the threads as needed"
(Much) more info in the above-mentioned PDF.
The scheduler in the kernel does do this if possible. In order to maximize the chances of a process waking up to warm caches, the scheduler always attempt to wake up the process in the core that was most recently running it. If that core is not available then it tries another core on the same package before it tries other packages (if there are many cpu packages).
However, the kernel has other real world priorities than just trying to keep a single process throughput as high as possible. It must also be fair to the other processes while keeping the current consumption low and as many cpu cores powered off as possible.
Because the kernel has to work with different workloads ranging from low-end not-so-smart phones with a battery to a supercomputer with it's own nuclear power plant, there are compromises to be made. There's also a lot of compile time and runtime configuration options you can use to tweak the kernel to your particular workload.
In short, in current CPU generations (x86_64 is what I work with) the PCIe controller is on the CPU die so if you have multiple CPUs you have multiple PCIe controllers and each of them control different PCIe sockets. You need to consult your motherboard to know which sockets will go with which CPU.
If an interrupt comes from a PCIe card it will trigger on the CPU that is attached to it, if the process that needs to handle the data is on the other CPU you need to transfer the data and the cache to the second CPU, it's a small effect but if you really care about performance and want to squeeze every nanosecond of latency on your work you should care about this.
You should start by considering in which slot to stick which card and if you have a multiple of PCIe cards you really want to balance them out.
It's also useful for any other PCIe card, sans the IB specific parts.
So personally, I have my desktop CPU (i7) running at performance setting. That was clearly noticeable for me for short bursts of activity. It's possible you could settle on doing that for just some of the cores (I don't know if that makes sense, since they're in a single package).
The I/O and Ram paging situation far dominates.
In fact, this seems like a Prisoner's Dilemma problem with the OS having programs cooperate by default. There's a chance you'll speed up your application by locking it to a core (defecting), but only if none of the other applications try the same thing.
SGI had this in Ultrix in 1991-1992ish (maybe earlier).
Why? Their market was graphics. If anything non-essential prevented the system from rendering a frame, an attempt was made to run that on another processor.
I remember a 4 processor reality engine^2 that had four processors. 1 processor was dedicated to graphics. another to networking. I don't recall what the other 2 were for.
What you're thinking of is different from task/processor/thread affinity.
Regarding SGI, I'm guessing you're thinking about the Onyx, which would have put it about 1993+. The R4x00s that were in the machine were not powerful enough to drive the RE2 and IR boards on a shared basis. The scheduler on many of these machines (Ultrix and IRIX both, along with Unicos, etc.) all pretty much sucked at the time, with AIX being a notable exception in some environments, especially running under VM.
You still see dedicated task processors with the z-series and maybe i as well today.
If you have 2-threaded code, it's also fun to try taskset onto one core with hyperthreading.