NUMA Deep Dive Part 1: From UMA to NUMA
frankdenneman.nl
frankdenneman.nl
Obviously you are focusing on x86, but it might be worth mentioning where many of these ideas developed. For example, SGI Origin hardware (http://www.sgidepot.co.uk/origin/isca.pdf).
As to your question on what CPU the kernel normally runs on? AFAIK, it normally starts at core 0, processor 0. You can see for yourself where processes run.
For example using the ps command.
It will print kernel threads in square brackets and it can also list cpu and core number behind a /.
Eg, example output of a linux VM with 2 processors.
$ ps -ax
PID TTY STAT TIME COMMAND
1 ? Ss 0:03 /sbin/init splash
2 ? S 0:00 [kthreadd]
3 ? S 0:00 [ksoftirqd/0]
5 ? S< 0:00 [kworker/0:0H]
7 ? S 0:15 [rcu_sched]
8 ? S 0:00 [rcu_bh]
9 ? S 0:00 [migration/0]
10 ? S 0:00 [watchdog/0]
11 ? S 0:00 [watchdog/1]
etcetera.. removed the rest for readability
The CPU scheduler then handles the rest to see on which thread/core to run a new process.This post might also help:
http://superuser.com/questions/389161/what-do-the-mean-in-ps...
Whenever user mode does a trap to supervisor mode it is running the top half of the kernel. An interrupt (device or timer) causes the core to run the bottom half of the kernel. Many systems used to only send these to cpu0, and it is typical that lots of network I/O would occur on cpu0.
In the PIII era, with a wide bus shared by all cores, most kernel data structures were not multithreaded, i.e. single run queue, great big kernel locks. There are now many per-core data structures, which can be in local memory. See http://www.tldp.org/LDP/tlk/kernel/kernel.html and for a recent discussion of scheduling: http://www.ece.ubc.ca/~sasha/papers/eurosys16-final29.pdf
I need to educate myself to understand the consequences of this decision, what the payoff would be to enabling NUMA, what my configs should look like, if there's an "easy way out", etc.