Rpath, or why lld doesn’t work on NixOS
matklad.github.io
matklad.github.io
That’s better since editing a binary will break its codesigning.
Suppose there are two installations of app. One is the system one, and one is locally installed by the user. The user overrides LD_LIBRARY_PATH when invoking the local app. Suppose that that app is used in such a way that it invokes the system-installed app; that could then find the wrong libraries due to the LD_LIBRARY_PATH being inherited.
A program must simply know where its exact pieces are, all by itself, without any external tricks that could influence more than just that program.
Search paths (all of them, including PATH) should be left to the user, for arranging the system; the user should be able to manipulate paths in arbitrary ways, yet the application shouldn't break as far as being able to locate and load its own pieces.
When developing ArchMac I had to do godawful hacks to bog-standard libs because whatever build system decided to hardcode a lib path (or forcefully strip one when it should be hardcoded, I've had to handle both) that I had to manipulate through various means including install_name_tool which is not that different from patchelf†.
This kind of issue was not macOS specific, it just turns out the various ways things were built happened to gracefully "work" on most Linux distros by sheer luck but they could have been equally broken.
† Not really a surprise when thinking about it, the concept of Nix derivations is not that different from the concept of Darwin bundles/frameworks (in terms of being a self-contained dependency package) so it's only natural similar issues, and thus approaches and tools to tackle them, emerged.
In fact, I am running the `evdev` example and I don't get any linker errors at all even when I change the linker to LLVM. I am using a nightly version of rust though.
[1] It's not - it's MPL
Pretty sure that runpath and rpath are distinct and have slightly distinct behavior. Can't fault you much for making the mistake, though. The two are not given names to be easily distinguished.
> Such directories are searched only to find those objects required by DT_NEEDED (direct dependencies) entries and do not apply to those objects' children, which must themselves have their own DT_RUNPATH entries. This is unlike DT_RPATH, which is applied to searches for all children in the dependency tree.
Search order:
glibc: rpath > LD_LIBRARY_PATH > runpath > ld.so.cache > default paths.
musl: LD_LIBRARY_PATH > rpath=runpath > default paths.
Search path inheritance:
glibc: rpaths are inherited: When exe depends on libx depends on liby, then liby first considers its own rpaths, then libx's rpaths, then exe's rpaths. HOWEVER if liby specifies runpath, it will not consider rpaths from parents.
musl: rpaths and runpaths are the same and always inherited.
I verified the glibc/musl sources when writing https://github.com/haampie/libtree
https://fzakaria.com/2022/03/14/shrinkwrap-taming-dynamic-sh...
"As this is NixOS, we are not going to barbarically install it globally"
basically you build your binary and set a rpath, let's say it is /usr/local/lib in your machine
mach-o binary stores: @rpath = /usr/local/lib libxyz = @rpath/libxyz.dylib.1
when I want to install those to /opt/something only thing I need to do is install_name_tool -add_rpath /opt/something
This will add search directories to binary itself. There are some DYLD_* environment variables too but I'm not sure about them... (Some are SIP protected by the way)
PS: It may invalidate signed binaries. Again, not tested such use cases.
You can try to do this with rpath of $ORIGIN. But then you'll probably still run into libc issues -.-
Not hard to see why a lot of games focus on Windows and let wine/proton handle Linux.
The idea of doing that for the DLL search in Windows could have been inspired by C.
It's the original fix!
However there is also the WinSxS alternative, adding a manifest file to the executable.
GNU/Linux doesn't have DLL hell only to the extent that there is an entire binary distro with maintainers beavering to keep all of the dependencies straight so that every program that needs a certain shared library is maintained to need the same version of it as any other program.
You will experience shared library hell as soon as you have your own binary application that is not in the upstream distro, and it happens to depend on one of the lesser libraries that do not do symbol versioning like openssl, libbz2 and whatnot.
I've dealt with this plenty in more than one Linux embedded dayjob. In one case, I hacked an elaborate library searching system around dlopen() into such a program.
Linux (g)libc is sort of equivalent to the Win32 API on Windows so you are not expected to ship your own version just like you don't ship your own ntdll, user32, etc. Since glibc has good backwards compatibility the onlye libc issue you will run into is having to compile against the oldest version you want to support.
does precisely that. @executable_path is evaluated at runtime.
Since nix stores everything at /nix/store, I didn't want to mess with relative paths...
But yeah, macOS apps can ship their own version of libraries thanks to this flexibility.
A library itself decides if it is relocatable or fixed. If it is fixed the MH_DYLIB records its install name as /path/to/binary (generally by setting DYLIB_INSTALL_NAME_BASE so xcodebuild will merge that with the library name automatically). The binary must be at that path. However this can (and often is) a symlink just like other systems use where /usr/lib/somelib.dylib -> /usr/lib/somelib.1.3.dylib so that minor version updates can be made without rebuilding programs.
If a library wants to be relocatable it specifies an install name of @rpath/binary.
At runtime dyld creates a "run path list". Every time it encounters a load command with an @rpath name it tries substituting paths from the run path list until it finds the library. The main binary along with any dependencies can add entries to the run path list. These can be absolute paths or relative paths anchored from @executable_path or @loader_path. The former being the main binary and the latter being the path to the binary itself (eg if the main app loads a plugin the plugin can reference dependencies relative to the main app or itself as needed).
You can push your own paths in the mix with DYLD_LIBRARY_PATH (searched first) or DYLD_FALLBACK_LIBRARY_PATH (searched last). Check "man dyld" and "man ld" if you want more details.
None of the above requires modifying binaries and so doesn't invalidate code signatures. If you want to use install_name_tool on binaries you build pass "-headerpad_max_install_names" to ld so it will pad out the load commands which makes it easier to edit them.
There have been a bunch of security vulnerabilities around the Windows strategy of auto-loading any dependency from the same directory as the binary so YMMV.
My issue is the last sentence:
So… .. turns out there’s more than one lld on NixOS. There’s pkgs.lld, the thing I have been using in the post. And then there’s pkgs.llvmPackages.bintools package, which also contains lld. And that version is actually wrapped into an rpath-setting shell script, the same way ld is.
That means that there isn't a problem. NixOS has fixed this, the system works. Except that you have to magically know which package you should be using. This is the sort of problem that I run into with Nix - it's hard to know the correct incantation.> Curious observation: dynamic linking on NixOS is not entirely dynamic. Because executables expect to find shared libraries in specific locations marked with hashes of the libraries themselves, it’s not possible to just upgrade .so on disk for all the binaries to pick it up.
Dynamic linking is not just about being able to upgrade without re-linking. Dynamic linking is not even primarily about that, not anymore, if it ever was.
Dynamic linking is more than anything about semantics that no one has bothered to add to static linking!
Static linking for C is stuck in the 1970s.
Dynamic linking for C makes C more like C+ -- a different language.
Specifically:
- with static linking symbol conflicts are a serious problem
- with dynamic linking symbol conflicts need not be a problem because with direct binding (Illumos) or versioned symbols (GNU), you get to resolve the bindings correctly at build-time and have them resolve correctly at run-time
- at build time you get to list just direct dependencies, and the linker does the rest -- compare to static linking, where you have to list all dependencies only in the final link-edit and then you must flatten the dependency list into some order, and then if there are conflicts, you lose.
For all those who keep harping on how static linking is better than dynamic linking, what I would suggest is that what must be done to make static linking not suck is to enrich .a files with the kinds of metadata that ELF adds to shared objects so we can get the same "list only direct dependencies" semantics when static linking as when dynamic linking. And I would note that libtool does this, just... very poorly.
What I would do to fix static linking:
- have `ld` add a .o to every .a that includes the `-L`/`-R`/-l` arguments given when constructing the .a (normally one does not do this when linking statically!)
- have `ld` look in every .a found when doing a final link-edit to recursively find its dependencies, and, most importantly,
- provide the same direct binding / versioned symbol semantics as in dynamic linking so that external symbols in dependents are resolved to the correct dependencies in the same way as in dynamic linking.
Notionally this is quite simple. But adding this to the various ld implementations would probably be rather a lot of work. Still, if people insist on static linking, this work should be done.
Or just use cmake which does that automatically
The problem is in the link-editors. That's where the problem must be fixed.
At least meson can parse those. Making every other build system in existence add support for those is definitely less work than changing the format of .a and will yield exactly the same result (not that it matters much, cmake being the standard c/c++ build system for years now)
Also .pc files are near inexistent on the most used desktop OS.
And that is bad. `-lfoo -lbar` loses important information. It is the linker-editor that needs to have this, and not just the build tooling layered above it. libtool, cmake -- it doesn't matter, they're all broken for static linking as long as static linking is broken in this way.
> Of course this information isn't stored in the .a files, but in cmake's FooConfig.cmake files - who cares as long as it works ?
If the only way to make it work is to adopt a particular build system, then no thanks. But again, it can't actually work -- it can't solve the problems that are in the link-editor itself.
The main app I work on links against
- Qt
- LLVM
- libclang
- ffmpeg
- my app which is itself ~40 libs
- a dozen others
which combined account for multiple hundreds of static libraries on Linux, Mac and Windows and things just work. I could maybe bog my head against theoretical link-time problems which I don't experience or make things better for my end-users and just target_link_libraries(myapp Qt::Core), what do you think is the reasonable option ?
But libtool is written in POSIX shell, it's not part of the linker, and it is a bit of a disaster.
The first one is to make the linker (ld) happy: -la will look for liba.so. The linker puts the SONAME (liba.so.x) in DT_NEEDED.
The second symlink's filename corresponds to the SONAME, so that the runtime linker (ld.so) can locate the library by SONAME in rpaths.
The third one is the actual library, which can be updated while keeping the same soname & same abi.
Now, it would be great if the linker had an option to not only copy the SONAME into DT_NEEDED, but also register the path in which the library was located as an rpath.
Cause the situation on Linux is absurd! You pass some flags -L and -l to the compiler/linker, the linker links something and nobody knows what. Then when you run your executable it has to locate this something again, and you can only pray that your libc and binutils/llvm agree on search paths & order. In most cases this does not work, and you must manually pass -Wl,-rpath,/some/path to add a search path. Nobody guarantees that what the linker links is what the runtime linker uses.
Of course there are many edge cases:
- linking during make without relinking during make install will make your executables register rpaths to build directories instead of install dirs
- sometimes you link to a stub lib that should not be used at runtime
But still, some guarantee that what you build with is what you run with would be a major user experience improvement for linux.
The issue you're missing is that the build directory is usually not the location of the final binary objects. The actual absolute path of libfoo.so may well be in /builds/runner/foo-package-4df78af0/build/prefix/lib/libfoo.so, which is unlikely to exist on anyone other than the CI's machine (and even on the CI machine itself for too much longer). The actual location will usually be /usr/lib64/libfoo.so, but the library that is linked against may well not be there at the time of linking (particularly in the case where a package is building both a library and an executable that depends on said library in the same package).
What you really want is for the relative path to the library to stored in the executable. Unless what you want is to actually use the globally-installed library and not one you're building at the same time. There's no single solution that fits every use case!
Usually indeed, for distro's that pretty much support one single version of every library. But this is no longer true for Nix, Spack, Gentoo Prefix and Guix, all these package managers/distro's have in common that there should be no default search paths where all libraries are dumped into.
How about `--copy-link-path-as-rpath` and `--copy-link-path-as-rpath-ignore=/build/dir`, so that ld continues to copy the soname to dt_needed, and registers rpath of non-build dirs. Then Nix, Spack, ... can simply use these flags in their linker wrapper.
My build system will have already resolved all these paths. It's very easy to interpolate these paths into the command to call the compiler.
Your complaints are reasonable, but gcc already does this. Here, I'll show you how:
> You pass some flags -L and -l to the compiler/linker, the linker links something and nobody knows what.
I agree, it would be really nice to be able to specific exact shared object paths instead of using -L and -l. Build systems typically already know the full paths to all the objects and the abstraction here is often unhelpful.
This could be remedied fairly easily by allowing (for example) -l to take an absolute path to an object rather than searching -L paths. But gcc already does this - you can just put the shared object on the command line directly like so:
Change this: gcc -lfoo bar.c
Into this: gcc bar.c /path/to/foo.so
The effect is the same, but more explicit. The foo.so.x object will be linked and added to the SONAME list.
> Nobody guarantees that what the linker links is what the runtime linker uses.
This part, however, is by design. We explicitly do NOT want what the linker links to be what the runtime uses. This is how we update shared objects between minor versions to fix bugs without reinstalling every binary on the entire system!
Letting libfoo.so.1 link to libfoo.so.1.x is a huge feature. Locking in an explicit minor version would defeat the entire purpose of dynamic linking.
My suggestion is to continue copying the SONAME into DT_NEEDED, and record the dir the lib was found in as an rpath.
I did not say it should fix the lib by filename.
Regarding runtime, we absolutely do not want to implicitly embed build paths. As others have said because it's unlikely we will put them in the same place, and it's unlikely we will build and run on the same systems.
The runtime library management system is going to have a structure for where to place libraries. It may be as simple as tossing them in /usr/lib, or it may be somthing where we have different paths for each application. We can't do this if the compiler implicitly dictates universal linker paths.
One aspect I think you may also be missing is that shared objects themselves have DT_NEEDED and rpaths. You would quickly run into very confusing conflicts between binaries built on different systems, or with different build environments.
It's hard to see a problem here, since adding an rpath is very easy. You appear to be asking for implicit, hidden behavior in the compiler which doesn't fit the vast majority of use cases.
I suspect that NixOS is playing with this in order to have a relocatable install: so that is to say, so that user can install NixOS in some subdirectory of a system running some existing distro. Any subdirectory, yet so that programs can find their libraries.
If I were in this predicament, rather than perpetrating hacks to patch the rpaths in binaries, I'd fix the dynamic linker to have a better way of locating shared libraries. The linker would determine the path from which the executable is being run, calculate the sysroot location dynamically, then look for libraries in that tree. E.g. /path/to/usr/bin/program would look under /path/to/usr/lib and related places.
A possibly nice hack would be to extend the meaning of the rpath variable. Give it a syntax, like say that if it starts with @, then the rest of it denotes a relative sysroot path fragment.
E.g. the program that gets installed as /path/to/usr/bin/program would be built with an rpath of "@/usr/bin". So then the dynamic linker sees the @ and does a sysroot calculation. First it strips off the basename to get just the directory part "/path/to/usr/bin". Then it sees, hey, the suffix of "/path/to/usr/bin" matches the "/usr/bin" in the rpath. The suffix is stripped to produce "/path/to" and that path is then used as the root for the library searching. Instead of searching literally in /lib or /usr/lib or whatnot, the "/path/to" part is prefixed to every search place to look in /path/to/lib and /path/to/usr/lib.
Patching binaries is very poor; it changes their cryptographic hash like SHA-256. You want your distro to be installing bit-exact stuff from the packages, and treating it as immutable.
Kinda the opposite: NixOS doesn’t really have a sysroot.
This sounds vaguely similar to dyld's @executable_path variable on macOS.
Binary patching only comes in when they're trying to get closed source binaries to run on NixOS and is fully managed by the packaging process to happen the same way on every system. These days, though, the approach of using filesystem namespacing to give packages a custom FHS-like view of the world seems to be growing more common instead of the patching.
That said, it's generally very difficult to build deploy-time relocatable code in Unix-land. The problem is that there's nothing like $ORIGIN for finding static assets, and all the autoconf tooling just makes it so easy to make all paths in object code by absolute paths that include the install $prefix/$bindir/$libdir/$sharedir/$statedir/$etcdir, etc.
Not that one cannot write deploy-time relocatable code -- I've done it plenty. But that it requires so much foreknowledge, intent, and know-how, that it just doesn't get done.
And for people saying "Buy you can't get security updates".
I would rather have a dynamically-linked binary that includes all the dependences in it, at which you can upgrade the dependencies with a tool over the binary, than the madness of shared libraries in system paths. (Well, you kinda get that with appimage and similar).
There are good reasons to use LXC and namespaces and shit, but mostly Docker is a workaround for how fucking stupid dynamic linking is as a default.
There are even use cases for .so, but as a default? Someone should be flogged. It was stupid when Sun pushed it with X in the 90s, and it’s stupider now.
Also, gold/mold are worth trying in addition to lld. Depends on your software.
Using Nix without relying on symbolic links would be ... challenging.
Excluding that, you can go the Alpine/Docker route, which isn't terrible but isn't great.
Or you can rebuild your box via Ansible or whatever every time you need a new AMI. This is actually a pretty reasonable solution.
I originally used Absible but switched to Nix. I found ansible to be too idiosyncratic and brittle to maintain. It also isn't inherently idempotent though it's supposed to be used that way.
The `glibc` people won't even make it work properly without being an `.so`. I appreciate that the GCC people are desperately trying to hold on to relevance in the face of a much better toolchain, but at what point does the FSF die a hero or live long enough to become the villain it was trying to conquer?
Jeremy is a bright guy and I have huge respect for him. And I know the `cat-v` folks troll a bit much, but there is some real substance to linked set of arguments [0]. And it very much ties out with my experience.
In the glory days at FB we statically linked everything, and it was amazing. I don't know this firsthand but I've heard it secondhand, and given that we plagiarized basically everything else from Google in those days, I tend to believe the claims that Google did the same.
“I tend to think the drawbacks of dynamic linking outweigh the advantages for many (most?) applications.” – John Carmack
Dynamic linking used to be a big optimization thing -- saved on disk space, saved on virtual-memory footprint; nowadays we hardly notice. Next, it meant you didn't need to rebuild everything when a library got a fix or backward-compatible improvement. Then, it became a way to get security patches into use quickly.
It is kind of impressive that we don't (often?) see dynamic linking itself used as an attack vector, aside from bugs in the libraries so linked.
Silently patching `.so` code actually obscures whether or not you have the relevant security fix to the relevant library. `libfoo.0 -> libfoo.0.1 -> libfoo -> 0.1.3` would be confusing enough even if vendors didn't routinely change the code out from underneath without moving the "version". If you're linking a `libfoo.a`, not only does the hash of that file change when it gets updated (or was failed to be updated), but the hash of your resulting binary that you choose to run either did or did not change. You don't need NixOS for that. No `LD_PRELOAD` crap can get in front of you because someone grabbed control of environment variables. There's like a zillion less things to go wrong from a security perspective. And even then, the security that you get from running a binary you built against someone who is already on your machine, and already has enough permissions to run it? You're in murky territory already. Keep them out of your box unless you're a cloud provider or something.
I'm no FS expert and I'm prepared to believe that symlinks might create problems for filesystem engineers, but you also sort of need them if you want atomic FS operations on POSIX. Something being a standard isn't a blank check to be bad obviously, but it's a lot easier to change whether or not `glibc` breaks on purpose when statically linked than to change a 30-year-old standard.
Dynamic linking is a nightmare for a number of reasons: it makes it murky and difficult to know what code is running, it makes effective text segments depend on environment variables, it further privileges superuser, it destroys the portability/backwards-compatibility that Linus has fought so hard to preserve in the kernel, it miseducates people about how virtual memory works by leading them to assume that it's some kind of performance win (it's not), it complicates the whole system by a ridiculous amount, it requires that you `readelf -d thing` to even know what crazy `rpath` shit is going on.
Symlinks have problems, I've been burned by them. But you can use them or not as you like, and the complexity is manageable. You show me a serious hacker, I'll show you someone who can get symlinks substantially right.
You show me someone who really, really deeply understands what the hell is going on with dynamic linking? I'll show you Ulrich Drepper and his weird agenda around suppressing LLVM.
If you code a program that does Posix file system operations, and may be used where security matters, congratulations, you are a security hole vector.
I appreciate that for a lot of desktop users they just want it to work and aren't terribly picky about which minor version of `libfoo` is required to get Firefox or Chrome to start.
But there are vendors and maintainers who are deeply concerned about such things when attempting to give the desktop user a seamless experience, and why a big mushy puddle of who-friggin-knows code makes that job any easier?
Who friggin knows. I think it's just inertia.
The fact of the matter is that everything from snaps to flatpaks to docker to this whole containerization craze is mainly dealing with the pick a card, any card outcome you get with a bunch of random `.so` in `/usr/lib/`.
Dynamic linking is a tool to improve system administration.