Ubuntu Looking at Applying Low-Latency Optimizations to Its Generic Kernel
phoronix.com
phoronix.com
> Ubuntu's low-latency kernel is mostly Kconfig configuration changes applied to their kernel build that they are now at least considering to make by default.
Ubuntu used to have an absolute monster of a hacked up franken-kernel compared to mainstream, but (very fortunately) over the years it's gotten much, much closer to upstream. It was honestly a bit of a nightmare trying to apply a set of custom patches on top of it (to improve a driver that was buggy for many years that I hacked in a way that made it work awesome but would never be accepted upstream because it violated FCC regulations). I'm an outside observer so pure speculation, but I have suspected that it was painful for the Canonical kernel team just as it was for me, and that they made a concerted effort to get away from the craziness. In related news, the people who backport kernel fixes to LTS versions they maintain for their distros (think Canonical and Red Hat, for example) are some of the most unsung heroes in the world. That work is brutal and takes a lot of dedication, and when you're doing your job perfectly nobody knows, but when you mess up everybody knows. That leads to a lot of thanklessness.
Now I want to know more! Was it a misbehaving wifi? Did it violate the spirit, or the letter of the regulations?
I've had the opposite problem, trying to make a card respect FCC regulations (ieee80211_regdom=US) while its firmware seems to have another country hardcoded.
I think it would be far better to use regdom=US instead of letting the card do whatever's ok in the country the card believes it's still operating in, yet it would also "violate" FCC regulations to let me write into its firmware that it should follow US rules by default from now on.
ebay + a worldwide market + "can't trust people to change the firmware" = problems
I wish there was more flexibility in the drivers to correct such misbehaviors, because I believe most people want to do the right things.
Surely it'd make way more sense to just have a locale configuration option, rather than only allowing the intersection of all radio emission rules worldwide?
I can easily run buffer sizes for 0.5ms latency on the rt patchset while running random twitch streams and games at the same time, but struggle to maintain 10ms on lowlatency without xruns with everything but the core audio apps shut down.
I've had no trouble plugging in random PCIe or USB devices and getting <6ms buffers. Everything works fine. After some years deeply enjoying JACK's incredible flexibility & latency, I eventually switched to PulseAudio with mostly stock config but much reduced buffers.
Now a days PipeWire is the best of both worlds. It'd be interesting to see if the top post still had such issues as they do under PipeWire, although I'm unsure what about pipewires scheduling would have better mechanistic sympathy.
Plenty of distros do offer -rt kernels. But -rt has been a massive set of patches for a long long time, and it's a different operating mode, and it has tradeoffs. Extreme latency sensitivity is a niche need, and giving it to everyone would have negative impacts. That's why it's not the default. Good news though, -rt's longkng effort to upstream appears to be nearing a successful conclusion, so realtime may be a Kconfig away in the future. We just have to solve printk() first, ha! https://lwn.net/Articles/951337/
Maybe pipewire's use of DMA-BUF lets it deliver to the audio device faster, or maybe jack has a slightly better event loop/thread architecture. (Just making up random hypothesis here; these aren't real claims.)
It's not going to deliver miracles; you're right that these two systems attempt very similar tasks. But there is some variance of implementation, and likely both of these systems have some small overhead. Maybe jack already delivered the absolute best responsiveness that could be possible in the kernel, had found every iota of optimization to make, and no new possibilities have emerged in the 20 years since. Or maybe jack already was p100 responsive to 0.00001ms, small enough to be effectively zero.
You're totally correct that the kernel is the limiting factor. But with audio we need p100, buffer underruns are never acceptable, and maybe one architecture let's us run stably with a slightly smaller buffer then the other architecture.
If you really are that confident in decades old JACK being best userland can possibly for p100, congrats on the long path of analysis that's lead you to that rightful confidence.
I hope other major distros follow suit.
For example, the old process has its data in the CPU's fastest on-die cache, the new one doesn't - so it's initially slower. And the same thing can happen when switching back.
For the lowest latency applications, there's also competition from "just dedicate a core to it full time" and "offload it to a dedicated microcontroller" lowering the demand to get it merged.
That's interesting, is there a place to read more about it? E.g. how to set it up on Linux and how much does it help with audio.
Assigning an entire CPU core to audio might be a bit much, unless you've got CPU cores to spare - the context in which I heard of this was high frequency trading, where the server really didn't need to do much except run this one application.
I work in trading, pretty much all of our performance sensitive apps use core isolation.
On windows this sort of comes out of the box because windows always defaults CORE0 to the OS and is mostly single threaded, but you can set CPU affinity on your processes in process: just avoid core0.
https://learn.microsoft.com/en-us/windows/win32/api/winbase/...
I've found a lot of drivers like to spread their interrupts all over the place, so it's not as simple as using cpu affinity (or pinning processes) outside core 0 :(
Check the specs or ask GPT ("on many separate on-die cache does the 12th gen intel Alder Lake P i7 1270 has?") and you'll see it's simpler to think about it as having 3 groups of cores with different features because:
- the 1st group has 4 performance cores each with their own L1 and L2: that's where I put the OS, with core 0 as usual and the other in NO_HZ, and the interrupts on power core 4 (here I'm using 0 and 4 to match what coretype reports P-CORE=0..7 that's due to hyperthreading)
- then the next 2 groups have 4 efficiency cores each, sharing the L2 (The L3 is shared all over, so it doesn't matter)
You could try to use my configuration as a base, and if it's not enough, you could try pinning the audio app on the power cores (obviously not 0 and 4 that are used, but like 1-3 and 5-7, which you could make tick if you prefer, you could also use more than 100 HZ depending on your needs)
It will take you some experimentation, but as a fellow audiophile I'd be curious about the results it'd get you!
The CK patches were always something that many of us eagerly applied at the time.
I care about latency and I use the "lowlatency" flavor of the Ubuntu kernel. It's nice.
Huh? As a Gentoo user who searched about all of this, once upon a time, I remember this being worse for latency, introduced as an Android patch for battery life gains. Also, what's the point of setting CONFIG_HZ=1000 with it?
Agreed, if you go "tickless", you want to go to the lowest value for what's still ticking: I have my own kernel configs, and I generally use HZ=100:
# zcat /proc/config.gz |grep HZ_
CONFIG_NO_HZ_COMMON=y
# CONFIG_HZ_PERIODIC is not set
# CONFIG_NO_HZ_IDLE is not set
CONFIG_NO_HZ_FULL=y
CONFIG_HZ_100=y
# CONFIG_HZ_250 is not set
# CONFIG_HZ_300 is not set
# CONFIG_HZ_1000 is not set
CONFIG_MACHZ_WDT=m
For Little.Big heterogeneous cores (like on the i7-1270P Alder Lake P), I suggest also thinking about how the cores are arranged to group the interrupts intelligently, and not just go by which cores are efficiency/power as /usr/local/bin/coretype simply reports:
P CORES: 0..7
E Cores: 8..19
You should check which cores are on a core.id with:
cat /proc/cpuinfo |egrep "physical|core.id|cache.size|processor" |grep -E "core.id|processor"
On my cpu: processor : 0 core id : 0 processor : 1 core id : 0 processor : 2 core id : 4 processor : 3 core id : 4 processor : 4 core id : 8 processor : 5 (....)
Therefore you could want to:
- Leave efficiency cores 8..19 as-is (nohz_full is not ideal and consumes power, leave the efficiency cores alone - they are made to be efficient!)
- Use power core 0 as normal
- Put all the other power cores 1-7 but 4 in NOHZ_FULL
- Put all IRQ and callbacks on power core cpu 4: for performance + a race to sleep
So on my i7-1270P, I use: `nohz_full=1-3,5-7 rcu_nocbs=0-3,5-7 irqaffinity=4`
In theory, you get the best of both worlds.
In practice, unfortunately, it's not perfect as the NVMe and WIFI drivers require extra care or they'll sprinkle interrupts everywhere which you can check with just cat /proc/interrupts
For NVMe, you can limit the queues for nocbs with `nvme.poll_queues=1 nvme.write_queues=1`
For wifi, you can try to limit its eagerness to have more interrupts with a simple script like `for irq in $(grep -v " ...: ......... 0 0 0 ......... 0 0 0 0 0 0 0 0 0 0 0" /proc/interrupts |sed -e 's/:.//g' -e 's/^ //g' -e 's/ .//g' |grep -v '^$') do [ -f /proc/irq/$irq/smp_affinity_list ] \ && echo 4 > /proc/irq/$irq/smp_affinity_list \ || echo skipping $irq done`
BTW, if you have more IO and you prefer having interrupts spread a little more, you can use both cores one of the avx512-less efficiency core.id 8 with irqaffinity=4-5
Then the cmdline becomes: cpu0_hotplug nohz_full=1-3,5-7 nr_running=1 rcu_nocbs=0-3,6-7 irqaffinity=4-5
EDIT: I added some clarifications, it's very late but I realized it required a little more context
NO_HZ config options are there to remove ticks if theres nothing in the process queue.
This just means "Check if >1 processes need CPU time from the kernel, if not, skip";
The drawback is that it can actually be worse for latency, as you say, because there is an kernel<->userspace buffer that has to be checked and this is pure overhead.
However, if your process is mostly userspace (and almost no kernel interaction) then you will see better performance, because your process won't have CPU stolen by the kernel if there's nothing to be processed.
Since it doesn't remove ticks entirely, (just aborts ticks early); setting a high value for your tickrate gives you the best odds of getting your kernel bits rid of as soon as possible to minimise latency when there are ticks.
I think they're trying to compromise betweeen having a low tickrate (so a process gets a lot of CPU time for each tick) and a high tickrate (so kernel calls don't take long before they see CPU time).
I think the combination needs to be tested with your workflow, they conflict and might make things less good for many workloads (heavy IO, data processing, compiling)