Excellent succinct breakdown of the xz mess, from an OpenBSD developer
marc.info
marc.info
Lots of dependencies -> bad. Reducing dependencies will sometimes necessitate open-coding things, which too can be bad.
But what's interesting here is that if you're trying to find a good place to insert a backdoor, all you need to do is analyze the dependency graphs of targets of interest looking for ones which might be vulnerable to compromise via social engineering or other vectors. We should do our own such analysis looking for places that need to be watched better or removed.
Discussion: https://news.ycombinator.com/item?id=39903685
kinda wish this was unpacked a bit more, why exactly is a service executable dynamically linking to a library without using any of its symbols or functions, because of systemd.
and a follow up if some openbsd folks can comment. over the years i've read about various unique security capabilities in openbsd, it seems natural to ask, what kernel or OS capabilities does openbsd provide to thwart the stage 2 efforts for this class of injection techniques?
> The libsystemd library provides functions that allow interacting with various interfaces provided by the systemd(1) service manager, as well as various other functions and constants useful for implementing services in general.
https://www.freedesktop.org/software/systemd/man/latest/libs...
E.g. the service readiness protocol is pretty trivial to implement yourself if you don't want to pull libsystemd as dep.
If I recall correctly based on what I've read about this — I believe from the original mailing list post that noticed the vulnerability — it's because under certain circumstances in order to enable certain functionality you might want sshd to be able to talk to systemd, so distros often patch sshd with code to do that. But obviously you need a library to implement actually speaking system's protocol, and as it happens, the easiest way to do that is to include the entirety of libsystemd, since it has functions for doing that, even though there are at least two other libraries that implement just the communication functionality and are actually designed for non-systemd programs to use. The problem with that idea being that libsystemd, being a whole standard library for all systemd-related functionality that is probably mostly designed for use by programs in the tightly integrated systemd family and as a reference imenentation, also includes a lot of other code, including code that has to deal with compression that depends on liblzma, even though all of that is never used by sshd, because it only uses the small subsection of the library it needs.
The important question is: if one of those libraries is used, and then something else pulls in `libsystemd`, will they conflict?
That is overstatement. The docs have basic self-contained example how to implement the notification without libraries, its 50 lines, majority of which is just error handling:
https://www.freedesktop.org/software/systemd/man/devel/sd_no...
https://salsa.debian.org/ssh-team/openssh/-/commit/cc5f37cb8...
systemd has several ways of starting programs and waiting until they're "ready" before starting other programs that depend on them: Type=oneshot, simple, exec, forking, dbus, notify, ...
A while back, several distro maintainers found problems with using Type=exec (?) and chose Type=notify instead. When sshd is ready, it notifies systemd. How do you notify systemd? You send a datagram to systemd's unix domain socket. That's about 10 lines of C code. But to make life even simpler, systemd's developers also provided the one-line sd_notify() call, which is in libsystemd.so. This library is so other programmers can easily integrate with systemd.
So the distro maintainers patched sshd to use the sd_notify() function from libsystemd.so
What else is in libsystemd.so? That's right, systemd also does logging. All the logging functions are in there, so user programs can do logging the systemd way. You can even _read_ logs, using the functions in libsystemd.so. For example, sd_journal_open_files().
By the way... systemd supports the environment variable SYSTEMD_JOURNAL_COMPRESS which can be LZ4, XZ or ZSTD, to allow systemd log files to be compressed.
So, if you're a client program, that needs to read systemd logs, you'll call sd_journal_open_files() in libsystemd.so, which may then need liblz4, liblzma or libzstd functions.
These compression libraries could be dynamically loaded, should sd_journal_open_files() need them - which is what https://github.com/systemd/systemd/pull/31550 submitted on the 29th February this year did. But clearly that's not in common use. No, right now, most libsystemd.so libraries have headers saying "you'll need to load liblz4.so, liblzma.so and libzstd before you can load me!", so liblzma.so gets loaded for the logging functions that sshd doesn't use, so the distro maintainers of sshd can add 1 line instead of 10 to notify systemd that sshd is ready.
The time window where sshd is started and not yet ready to receive connections is short, and clients will have a connection timeout orders of magnitude larger. The notify functionality is more relevant for things like Java middleware processes and clients that lack the functionality to poll and wait. Under most situations none of this is relevant for sshd. These patches solve a problem very few people have.
If you really have this problem, the systemd readiness is far from enough to solve the problem. The readiness is sent too early and there could still be permission problems that would cause sshd to be ready but the connection to fail. Even more relevant is local firewall rules that are completely out of scope for a readiness check!
Polling for readiness is the only robust way.
What makes you say that? The readiness notification is sent after the sshd has opened the listen socket, it literally is accepting connections at that point.
Keep in mind that this is only a problem in special situations, almost no regular Linux servers carry any services that care about being able to establish ssh connections, that cannot reconnect and use appropriate socket options. For other services than ssh, such as databases, this is much more common. And for those, it is not enough that the server has opened a listening socket.
When building distributed systems, this is something you need to think about. Not so such with local systems. But the same principles apply. And the only robust way is to poll for readiness. Signalling readiness is both complex, when the dependency chains are non-trivial, and prone to error, when readiness signals arrive out of order or are dropped or fail for some reason. This could be because of operations failure but make for hard to debug cases. Dependency chains that mysteriously stop because of permission problems with out of band traffic is both classic and unnecessary problem.
All of this complexity go away when polling. This is why you should adopt this design in the somewhat unlikely case you have clients that depend on being able to make ssh connections.
A lot of things (e.g. OpenSSH, Apache) don't use it natively, you won't tend to see it in the official sources of non-Linux specific software. Package maintainers who choose to foist^H^H^H^H^Huse systemd add it to those packages. It's quite fun working out why Apache hangs when you port a config and don't realise it needs mod_systemd loaded to stop systemd stamping on it.
There's an extra security wrinkle too: lazy symbol resolution is not the done thing now, it's preferred that symbols are resolved at loading time so important chunks of memory can be protected from other types of attacks (more clarity in one of the preceding articles: https://research.swtch.com/xz-script ).
The backdoored liblzma relies on a misfeature of glibc called ifuncs, where a library can override a function by calling a special init function in the library. This is so for instance if you have a version of a function optimized for AVX512, one ofor AVX2 and one not optimized for those at all, the init function would check which features the CPU supports and picks the best one. Seemingly ifuncs doesn't check the function being overriden is in the same library, and it was replacing OpenSSH's RSA auth function.
OpenSSH added a clean-room libsystemd-free implementation of sd_notify() after this fiasco, in the hope Linux distros will stop linking against libsystemd, this will appear in OpenSSH 9.8:
> The stage 0 shell snippet looks at first glance like a plausible part of > the poorly readable autoconf/automake tooling.
In fact, malice and incompetence are not necessarily mutually exclusive.
This very incident shows several instances where "Jia Tan" is being arguably incompetent, in addition to being clearly malicious: unintended breakage by adding extra space between "return" and "is_arch_extension_supported"; several redundant checks for `uname` == "Linux"; botched payload, so "test files" had to be replaced, with pretty fishy explanation; rather inefficient/slow GOT parsing, list goes on...
Even if this weren't the case, your claim would remain specious. "Establish intent" and "justify attribution to malice" are the same thing said two ways.
If there was one commit with a stray period that disabled the sandbox, sure, might be a typo.
But the other stuff where they're unpacking specific byte ranges from multiple places in the test files, that's not an accident.
Viewed as a whole, it would be extremely hard to view this as "adequately explained by neglect, ignorance or incompetence"
I would certainly attribute that to a typo if I was reviewing the code.
I believe the point is that given this context, it could/should not be construed as a typo.
Or is this just an argument over the use of the word "typo" instead of "mistake"? That wouldn't be very useful, IMO.
Am I correct in concluding that this backdoor might have been picked up sooner if the practice of shipping tarballs of source releases was dropped in favour of running builds straight off a repo tag?
I've setup and worked with many build pipelines in the .NET world over nearly 20 years and we manage fine without the concept of a source release tarball, but I realise C dev is more complex so I'm probably missing something?
I don't think so. The git repo did have everything needed. The release tarball thing was just an added benefit to the attackers, but not critical. IMO.
> I've setup and worked with many build pipelines in the .NET world over nearly 20 years and we manage fine without the concept of a source release tarball, but I realise C dev is more complex so I'm probably missing something?
.Net, but, portable .Net?
The issue isn't the language but what things you need to detect presence of in the target host when building, and if you're doing systems programming you're likely to need to detect many such things, and if you're doing more high level programming (e.g., an HTTP app that uses a DB) then you're not likely to need to detect many such things.
The real issue is the variance from standards (e.g., POSIX) and the fact that the standards (e.g., POSIX) are so far behind because they really just standardize the lowest common denominator.
What I mean is that Debian and Fedora presumably download the tarball and build from there, rather than cloning a repo and building from the repo.
If they did the latter, it would have been much more difficult to manipulate the build process to inject the backdoor that lived inside the testcase binaries, because that change to the build scripts would be out in the open.
But I could be mistaken...
build-to-host.m4 is not checked into the repo. Without the modified build-to-host.m4, there is nothing to kick off the exploit.
The actual final product you get is when publishing (building) the application itself, similar to Rust the way crates participate in the build process.
Nuget packages still can and do ship pre-compiled native dependencies but ideally you want to either provide your own (think sourcing libsodium separately and just getting the bindings from nuget) or instrument the corresponding .NET project to build a native dependency alongside it (very easy with <Exec Command="..." /> properties in csproj).
A community ISP. Last changed 02.05.2001, in case you missed it. ;) Millennium blues. I wasn't aware about the possibilities back then, but it's just before my time.
...which sums up the situation quite nicely.
This was a pretty bad attack on a certain ecosystem. The ecosystem will recover, or not, regardless of your feelings. Unless you're in a position to truly make a difference, just sit back and enjoy the ride...