Huh? As a Gentoo user who searched about all of this, once upon a time, I remember this being worse for latency, introduced as an Android patch for battery life gains. Also, what's the point of setting CONFIG_HZ=1000 with it?
Huh? As a Gentoo user who searched about all of this, once upon a time, I remember this being worse for latency, introduced as an Android patch for battery life gains. Also, what's the point of setting CONFIG_HZ=1000 with it?
NO_HZ config options are there to remove ticks if theres nothing in the process queue.
This just means "Check if >1 processes need CPU time from the kernel, if not, skip";
The drawback is that it can actually be worse for latency, as you say, because there is an kernel<->userspace buffer that has to be checked and this is pure overhead.
However, if your process is mostly userspace (and almost no kernel interaction) then you will see better performance, because your process won't have CPU stolen by the kernel if there's nothing to be processed.
Since it doesn't remove ticks entirely, (just aborts ticks early); setting a high value for your tickrate gives you the best odds of getting your kernel bits rid of as soon as possible to minimise latency when there are ticks.
I think they're trying to compromise betweeen having a low tickrate (so a process gets a lot of CPU time for each tick) and a high tickrate (so kernel calls don't take long before they see CPU time).
I think the combination needs to be tested with your workflow, they conflict and might make things less good for many workloads (heavy IO, data processing, compiling)
Agreed, if you go "tickless", you want to go to the lowest value for what's still ticking: I have my own kernel configs, and I generally use HZ=100:
# zcat /proc/config.gz |grep HZ_
CONFIG_NO_HZ_COMMON=y
# CONFIG_HZ_PERIODIC is not set
# CONFIG_NO_HZ_IDLE is not set
CONFIG_NO_HZ_FULL=y
CONFIG_HZ_100=y
# CONFIG_HZ_250 is not set
# CONFIG_HZ_300 is not set
# CONFIG_HZ_1000 is not set
CONFIG_MACHZ_WDT=m
For Little.Big heterogeneous cores (like on the i7-1270P Alder Lake P), I suggest also thinking about how the cores are arranged to group the interrupts intelligently, and not just go by which cores are efficiency/power as /usr/local/bin/coretype simply reports:
P CORES: 0..7
E Cores: 8..19
You should check which cores are on a core.id with:
cat /proc/cpuinfo |egrep "physical|core.id|cache.size|processor" |grep -E "core.id|processor"
On my cpu: processor : 0 core id : 0 processor : 1 core id : 0 processor : 2 core id : 4 processor : 3 core id : 4 processor : 4 core id : 8 processor : 5 (....)
Therefore you could want to:
- Leave efficiency cores 8..19 as-is (nohz_full is not ideal and consumes power, leave the efficiency cores alone - they are made to be efficient!)
- Use power core 0 as normal
- Put all the other power cores 1-7 but 4 in NOHZ_FULL
- Put all IRQ and callbacks on power core cpu 4: for performance + a race to sleep
So on my i7-1270P, I use: `nohz_full=1-3,5-7 rcu_nocbs=0-3,5-7 irqaffinity=4`
In theory, you get the best of both worlds.
In practice, unfortunately, it's not perfect as the NVMe and WIFI drivers require extra care or they'll sprinkle interrupts everywhere which you can check with just cat /proc/interrupts
For NVMe, you can limit the queues for nocbs with `nvme.poll_queues=1 nvme.write_queues=1`
For wifi, you can try to limit its eagerness to have more interrupts with a simple script like `for irq in $(grep -v " ...: ......... 0 0 0 ......... 0 0 0 0 0 0 0 0 0 0 0" /proc/interrupts |sed -e 's/:.//g' -e 's/^ //g' -e 's/ .//g' |grep -v '^$') do [ -f /proc/irq/$irq/smp_affinity_list ] \ && echo 4 > /proc/irq/$irq/smp_affinity_list \ || echo skipping $irq done`
BTW, if you have more IO and you prefer having interrupts spread a little more, you can use both cores one of the avx512-less efficiency core.id 8 with irqaffinity=4-5
Then the cmdline becomes: cpu0_hotplug nohz_full=1-3,5-7 nr_running=1 rcu_nocbs=0-3,6-7 irqaffinity=4-5
EDIT: I added some clarifications, it's very late but I realized it required a little more context