splice() and the ghost of set_fs()
lwn.net
lwn.net
We were making use of `/proc/PID/pagemap`, which is a kernel-generated file that previously would show the physical addresses of all the pages in a process's address space. Unfortunately, with the Rowhammer exploit, exposing this information - even for one's own processes - to unprivileged users went from being harmless to a security risk.
The first we saw of the change was when newer kernels started reporting zeros for all physical addresses, unless we ran as root. We raised this the LKML, explaining that we'd been relying on this feature to implement a somewhat esoteric optimisation.
Linus replied very helpfully - the security fix trumped userspace compatibility but he could see a secure way of getting us the information we really needed, given the technique we'd described. He invited us to submit a kernel patch and gave a few hints about potential gotchas.
I did work one up but the kernel community actually jumped on it as an opportunity to do more cleanup, so I ended up just signing off on the patch they produced. It was all a remarkably smooth and efficient process.
[1] Time travel debugging - http://undo.io
TIL
All stories about kernels describe it as an absolute that userspace is never broken, but that actually makes sense
$ ls -l /proc/$$/stack
-r-------- 1 tanel tanel 0 May 26 21:52 /proc/967141/stack
$
$ cat /proc/$$/stack
cat: /proc/967141/stack: Permission denied
$
$ sudo cat /proc/$$/stack
[<0>] do_wait+0x1c3/0x230
[<0>] kernel_wait4+0xaf/0x150
[<0>] __do_sys_wait4+0x85/0x90
[<0>] __x64_sys_wait4+0x1e/0x20
[<0>] do_syscall_64+0x49/0xc0
[<0>] entry_SYSCALL_64_after_hwframe+0x44/0xa9
Edit: Adding one more comment - my impression has been that the "no userland-visible changes" promise applies to system calls - how procfs presents data as human-readable text in the /proc files has changed every now and then before (I recall the sar command showing wrong numbers after a kernel update, for example).e.g. we've seen the format of signal stack frames change in the past but nobody relies on that layout so it's OK.
/proc is another one along those lines - technically somebody could probably complain if it changes but if nobody shouts then it will just drift over time.
If you go here: https://undo.io/udb-form/
Then you do have to fill out a contact details form but you will be able to download a trial of UDB, our interactive debugger. That gets you the Time Travel functionality.
(if you're a VSCode user then you can download the extension https://marketplace.visualstudio.com/items?itemName=Undo.udb and start from there instead)
If you want the full LiveRecorder experience - additional tool, library, etc then you do have to request a demo. https://undo.io/about-us/contact/request-demo/ - Mention that you had exchanged messages with me (Mark Williamson - Architect @ Undo) and I can help from my side if you have any technical issues.
(edit: s/issues/technical issues/)
Such a blind/naive compatibility shim, would likely be much slower than whatever hand-written slow path the developer has in place. If the whole point of a call is to be a fast/efficient version of something else, then the call is breaking its semantics if it isn't more fast/efficient than the alternative.
Because of this, it's better to just make autoconf et al detect splice(2) as absent for the given use-case (and fall back to the hand-rolled portable slow path), rather than relying on a naive kernel shim.
Not only is the autoconf solution not fixing the problem, it’s placing a massive undue burden on developers. Linus has been inconsistent here. Telemetry would have been helpful here in aiding this work (ie support it with the slow path but report the event so that distro maintainers could provide feedback on broken paths). Once you think you’ve eliminated the long tail of issues, then remove and see if anything remains broken that telemetry didn’t catch.
So if splice(2) no longer works when the source is e.g. /dev/random, then autoconf could attempt to compile+run a program that splice(2)s from /dev/random to a known-working sink, and see whether that program SIGSEGVs or not when run; and use that to decide whether to allow an --enable-splice configure flag to be passed, vs. bailing out if such a flag is passed.
Of course, there's the implicit assumption that if you're passing such a flag, the build environment is going to be the deploy environment; or at least, the build environment's feature-set will be a subset of the deploy environment's, such that the build environment's runtime features can be used as a conservative underestimate of the deploy environment's runtime features.
This "build is a subset of deploy" is usually a sensible assumption. Build environments are controllable, while deploy environments are arbitrary; so anyone who wants to make a build that works on many different deploy environments, can just set up their build environment to have the "lowest common denominator" of the features of the systems they want to target.
(Compare and contrast: microarchitectural optimizations. Same story.)
And then the next day you, or your user, update your kernel, the result is out of date and your program crashes. That's just not how autoconf (or cmake or meson for that matter) are used.
> running autoconf checks inside a target-machine emulator
Run tests are quite rare. Older autoconf used to use printf to detect the size or alignment of a type, but for 15-20 years it has instead been doing binary search (basically "guess the number") so that only compilation tests are needed instead.
(I am a former autoconf developer and GCC build system maintainer).
That definitely broke my assumption that kernel updates were well vetted for regressions.
And this article is an example of where the decision was made to break user space.
That seems like the kind of things that is easier said than done.
> That definitely broke my assumption that kernel updates were well vetted for regressions.
Shouldn't that be evidence for the contrary? One break in a decade or two of use isn't so bad. It's actually pretty good.
I expect some more users of Proxmox given Broadcom taking over VMware (e.g. I'd like to merge away from ESXi to Proxmox as I don't trust Broadcom). Hopefully it does the product and community well.
The kernel is an extremely complex piece of code, the developers can't be asked to test every kernel release against every piece of userspace software on every hardware configuration. That's one of the reasons that new code and significant changes require a bunch of mailing list discussion, reviews, and signing-off.
Also, Proxmox ships with its own supported kernel, it shouldn't be a big surprise that you ran into issues while straying off the beaten path.
With no regression tests, nearly no unit tests, and lots of bits of supported hardware that nobody on the Dev team uses day to day...
It's a miracle that quality is as high as it is.
I would argue that it's also a failure by the distro; unless there was something special about your exact setup that would have made the bug not show up in testing, I would argue that proxmox should have been testing updates to catch that kind of problem before users noticed.
Maybe this? https://forum.proxmox.com/threads/latest-update-5-13-19-4-pv...
On a well done "modern suspend/suspend to idle", I use disconnected modern sleep + Windows Media Player on Windows 11 to listen to songs on my bluetooth headphones for a cost of about 1% of the battery per hour (as measured and plotted with powercfg) which can come down to about half of it, 0.5%/h when not using Bluetooth.
I wouldn't want Bluetooth to prevent sleep (a 1% drop per hour is better than not sleeping!)
I also wouldn't want sleep to prevent me from using my Bluetooth headset (the difference between 0.5% and 1% is significant, but it doesn't matter much in practice if my computer can be usable in the morning)
This is one of the rare examples of "modern suspend" delivering on its promises, and blowing the good old ACPI S3 away: I never saw a drop of battery <5% on S3 suspend-to-ram unless it also involved S4 in a hybrid sleep of "ACPI S3 suspend-to-ram then S4 suspend-to-disk after a while or when I run out of power whichever comes first"!
The linux kernel does take bug reports: https://docs.kernel.org/admin-guide/reporting-issues.html
However, that bug probably isn't specific enough as you've described it, unless you can find the commit causing it (such as via a git bisect https://docs.kernel.org/admin-guide/bug-bisect.html), or come up with a clearer repro.
Alternatively, if you're seeing the issue on a distro-maintained kernel (such as on fedora/ubuntu/debian with their kernel package), reporting the issue to the distro maintainers may be more appropriate.
[1]: https://github.com/torvalds/linux/commit/dfbba2518aac4204203...
Long story short: there are a lot of things in your laptop generating interrupts.
Some of them you want to ignore, because it would cause the behavior you describe (preventing sleep)
Some of them you really want to listen to closely, because if it's an interrupt generated by say brushing on the power button, not listening to it means not waking up from sleep (traditional example: GPE96 on dells, cf for example https://bugzilla.kernel.org/show_bug.cgi?id=102281) until a longer press on the button generates an ACPI event or a powerup.
You can configure or finetune that with /proc/acpi/wakeup which hopefully will give you more context as to what other people have explained here.
(with_resource is a generic example. Search for such capabilities in your library.)
Computers exist to automate. Let the computer remember to close resources upon scope end.
Go's defer is also helpful in this regard.
set_fs(KERN)
*function_pointer()
set_fs(USER)
There did it before, so they knew about that option, and if they didn't do it, they must have had a reason.The issue with forgetting function calls like this are the edge cases, and the kernel has a lot more of those.
[1] https://gcc.gnu.org/onlinedocs/gcc/Common-Variable-Attribute...
This seems entirely successful to me and retired crappy and insecure behavior.
The idea that you can never break anything means that some historical mistakes can never get fixed.
The someone complains about how the language / operating system is a bolted on mess of nonsense and people go chasing after the next hot thing which has the benefit of just being newer and being able to get away with breaking changes because people on the bleeding edge put up with it.
The linux kernel is probably striking the right balance here, and I doubt that this one breakage is a sign of the decay and downfall of the Empire.
edit: checked the current implementation in go, and it's only doing that for unix and tcp sockets, so it'll be fine.
Or did you mean unlocking an encrypted device after unlocking and booting on the primary device? What does systemd have to do with that? Unless you want to type in the password manually every boot, you store the password (for the secondary device) on the primary device, add the password file to /etc/crypttab and be done with it, right?
Sounds like they should leave this change for 6.0 then.
They can just change the max. It's not that big of a deal.
I used to complain that they didn't have recurring subscriptions, but I believe they have fixed that! Just sign up yearly, at the worst you are a patron to good journalism that is not really happening elsewhere.
Just reading an article or two a week will give you loads of insight into what is going on in the tools you have everyday