Anyway that's not why we have shared libraries. They were invented so that memory managers could load the same library once, then map it into the address space for every process which used it. This kind of savings mattered back when RAM was expensive and executable code represented a meaningful share of the total memory a program might use. Of course that was a very long time ago now.
To the extent that shared libraries economize on memory, or are claimed to, it's when a lot of different processes link the same version of the same library, e.g. your `libc` (or the original motivating example of X). The tradeoff here is that the kernel has an atomic unit of bringing data into RAM (AKA "core" for the old-timers), which is usually a 4kb page, and shared libraries are indivisible: you use `memcpy` and you get it and whatever else is in the 4kb chunk somebody else's linker put it in. Not a bad deal for `memcpy` because it's in everything, but a very bad deal for most of the symbols in e.g. `glibc`.
The counter-point is that with static linking, the linker brings in only the symbols you use. You do repeat those symbols for each `ELF`, but not for each process. This is often (usually?) a better deal.
A little poking around on my system to kind of illustrate this a bit more viscerally: https://pastebin.com/0zpVqA0r (apologies for the clumsy editing, it would be too long if I left in every symbol from e.g. `glibc`).
Admittedly, that's a headless machine, and some big clunky Gnome/Xorg thing might tell a different story, but it would have to be 100-1000x worse to be relevant in 2022, and almost certainly wouldn't offset the (heinous) costs. And while this is speculation, I think it would be negative worse, because there just aren't that many distinct `ELF`'s loaded on any typical system even with a desktop environment. Chromium or whatever is going to be like, all your RAM even on KDE.
There are couple of great CppCon talks that really go deep here (including one by Matt Godbolt!): https://www.youtube.com/watch?v=dOfucXtyEsU, https://www.youtube.com/watch?v=xVT1y0xWgww.
Don't be put off by the fact it's CppCon, all your GNU stuff is the same story.
This personal assertion has no bearing in the real world. Please do shine some light on what exactly do you interpret as "expensive", and in the process explain why do you feel that basic software design principles like modularity and the ability to fix security vulnerabilities in the whole system by updating a single shared library is "trouble-prone".
I mean, the whole world never saw a real world problem, not the UNIX comunity nor Microsoft and it's Windows design nor Apple and it's macOS design not anyone ever, but somehow here you are claiming the opposite.
"DLL hell" [1] was absolutely a thing. It was solved only by everyone shipping the world in their application directories.
A single product can have a modular design, but modularity across unrelated products designed by different designers is a different concept that should have a different name.
In my opinion it’s just not something that really works. This is similar to the “ability” of some software to have multiple pluggable database engines. This always yields a product that can use only the lowest common denominator… badly.
But now imagine having to coordinate the sharing of database engines across multiple relying products! So upgrading Oracle would break A, B and fix C, but switching A and B to MySQL would require and update to MySQL that would break D, E and F.
That kind of thing happens all the time and people run away screaming rather than deal with it.
Shared libraries are the same problem.
And nearly all libraries eventually become "security-oriented libraries", given enough time. Everything is a hacking target eventually.
The amount of frustration dynamic linking causes is just not worth it in today's day.
I think that calculus looks very different 20, 10, maybe even 5 years ago, when getting packages rebuilt and distributed from a zillion different vendors was much harder.
Docker/containers do far more than just bundle shared libraries. We use a container for almost all of our CI builds. Our agents are container images with the toolchains, etc installed. It also means it's trivially easy to reproduce the build environment for those occasional wtf moments where someone has an old version of <insert tool here>.
Tools aren't the only part of the environment though. Setting the JAVA_HOME environment variable for a build can dramatically change the output, for example. Path ordering might affect what version of a tool is launched.
They also mean that our developers don't need to run centos 7, they can develop on windows or Linux, and test on the target platform locally from their own machine.
The nice thing about fully statically linked executables can just be copied between machines and ran. If Java worked this way, the compiler would produce a (large) binary file with no runtime dependancies, and you could just copy that between computers.
Rust and Go don’t give you a choice - they only have this mode of compilation.
Docker forces all software to work this way by hiding the host OS’s filesystem from the program. It also provides an easy to use distribution system for downloading executables. And it ships with a Linux VM on windows and macos.
But none of that stuff is magic. You could just run a VM on macos or windows directly, statically link your Linux binaries and ship them with scp or curl.
The killer feature of docker is that it makes anything into a statically linked executable. You can just do that with most compilers directly.
> But none of that stuff is magic.
I never said it was, I simply said that docker does more than statically link binaries.
>You could just run a VM on macos or windows directly, statically link your Linux binaries and ship them with scp or curl.
https://news.ycombinator.com/item?id=9224
> The killer feature of docker is that it makes anything into a statically linked executable.
The killer feature of docker is it gives you an immutable*, versionable environment that handles more than just scp'ing binaries.
The benefit docker provides here is giving users a simple, reproducible, downloadable environment with no external dependencies.
I'm not saying Docker is useless in the current build environment. On the contrary - Having a convenient way to sandbox, ship and execute arbitrary execution environments and statically linked executables is really useful.
I think this is such a great idea we should make debian work closer to this model. Nix is a fantastic step in this direction and it'd be great if more distributions followed suit.
Debian and friends fill your hard drive with of weird, non-reproducible versions of various dynamically linked libraries. The fact I can't take a binary on my x86_64 linux machine and run it on your x86_64 linux machine is an embarrassment. Moving toward static linking and fully reproducible builds would be a tremendous help here.
Docker feels like a kludgy hack. I'd love a world where docker isn't necessary because its best ideas have been lowered into the OS itself.
It's plenty common for binary distributions of software to come as tgz files. That's usually including a bunch of .so files.
Yes, and compatibility across distros / releases is very hit-and-miss, because they usually ship certain libraries and rely on the OS for others.
I’m not sure whether it’s possible to ship every library without running into ODR issues, but at that point, you’d loose all of the advantages of dynamic linking anyway.
Static linking is not a new thing. It was there since a beginning. Dynamic linking is a new thing compared to it.
The very opposite. You can do all the sandboxing you want without needing docker and you will avoid docker's big attack surface.
Docker has cleared become popular because it drastically eases application deployment and cleanup.
Static linking also solves a huge number of these problems, and in a generally better way.
Docker has overhead, an associated attack surface, bulky artifices, etc but it yields very nice benefits for workload isolation (env configs, storage mounts, etc) as well as cross OS/arch builds.
Is there a solution that combines the best of these options? A super-slim docker for statically linked apps that combines the ease of deployment with a much smaller footprint?
I feel like this could be a simple shell script, at least for 99% of real-world use cases.
Once you bundle multiple libraries into blobs you always have the same security nightmare:
https://www.researchgate.net/publication/315468931_A_Study_o...
(and on top of that docker adds a big attack surface)
One headache over another.
What’s to guarantee applications haven’t statically linked a library (in whatever version) rather than linked a “system” version of a library?
With the dynamic one you have both problems…
> What’s to guarantee applications haven’t statically linked a library (in whatever version) rather than linked a “system” version of a library?
It's easy to see which commands are given to the compiler and if there is a copied library if you have sources.
If we made all builds reproducible (like oasis) we could setup an automated trusted system that does builds verifiably automatically (like what Nix has).
Doing automated, local patching and rebuilding is far from enough. You need experts to correctly backport patches and test them thoroughly.
Oh, just like Gentoo, Slackware and LFS. ;)
Seriously, maybe we shouldn’t ignore history and learn why things are like this today.
It’s not just a simple switch you flick and recompile.
Do you even understand that you are advocating in favour of perpetually vulnerable systems?
Let's not touch the fact that there are vague and fantastic claims about shared libraries being this never ending sources of problems (which no one actually experiences) that when asked to elaborate the best answer is wand waving and hypotheticals and non-sequiturs.
With this fact out if the way, what is your plan to update all downstream dependencies when a vulnerability is identified?
Linux has never been anywhere close to resistant to privilege escalation bugs.
But statically linked systems are perpetually vulnerable by design, aren't they?
I mean, what's your plan to fix a vulnerable dependency once a patch is available?
Are you tracking each and every dependency that directly or indirectly consumes the vulnerable library? Are you hoping to have access to their source code, regardless of where they came from, and rebuild all libraries and applications?
Because with shared libraries, all it takes to patch a vulnerable library is to update just the one lib, and all downstream consumers are safe.
What's your solution for this problem? Do you have any at all?
Do you actually have an answer any of the points I presented you, or are you going to continue desperately trying to put up strawmen?
If any of your magical static lib promises had at bearing in the real world, you wouldn't have such a hard time trying to come up with any technical justification for them.
Unfortunately, this is very far from being true in the real world.
And even if all libraries did maintain 100% ABI compatibility forever, and even if there were never any compiler or processor bugs that needed to be mitigated by recompiling applications, dynamic linking would still add runtime complexity and thus surface area for security vulnerabilities.
> Are you tracking each and every dependency that directly or indirectly consumes the vulnerable library? Are you hoping to have access to their source code, regardless of where they came from, and rebuild all libraries and applications?
Yes? I mean, are you just throwing random untrusted binaries onto your servers, not keeping track of what their dependencies are, and hoping to God that you will never ever need to recompile them? (And if your explanation for that is "I use proprietary software from vendors who refuse to share source code", have you chosen incompetent vendors who cannot respond to security vulnerabilities in a timely fashion?)
1. For each application and library in the system, maintain a copy of the source code and the scripts necessary to rebuild it.
2. Keep track of the dependency tree of each application and library in the system.
3. If there's a problem with a dependency, update it and rebuild all dependent applications.
4. If I'm using binaries from a vendor (whether that's a Linux distribution or a proprietary software company), the vendor needs to be responsible for (1) to (3).
In practice, your package manager will generally implement almost all of (1) to (3) for you. (Newfangled systems like NPM or Cargo usually do it using lockfiles and automated tools like Dependabot.)
If you're using dynamic linking, guess what? You still have to do these four things, because even with dynamic linking, there are still problems that can only be solved by recompiling your software. Even if you think reasons like ABI incompatibility or processor bugs aren't compelling, there might be, y'know, bugs in the applications themselves that you've gotta patch.
I'm not even going to bother pointing how unfeasible all your other points are. I'm just going to point this Fack: you do understand this does not work and never worked at any point in time, don't you?
Do you understand the explicit reference to perpetually vulnerable systems? Do you realize where it cames from?
You only have the power to rebuild the packages you own personally, and even so static libs offer zero ways to keep track of which version went into which build.
Once your fantastic panacea starts to rely on your idea to force third parties to follow your personal orders to make new releases under your own personal terms, you should be very aware that you will not get your wish. You'll instead just keep on using the same vulnerability-riddled release.
Do you understand the importance and value of static libs? You do not need four-point authoritarian and deeply impractical and unfeasible plans to keep your system safe. With shared libraries you just patch the one lib, and all your system is safe.
And again what tradeoff do you want to achieve for this? Nothing?
Some libraries managed to preserve API contracts throughout the years just fine.
Even in C++, a language which does not establish a standard ABI, has frameworks which have not been breaking the ABI throughout multiple releases, such as Qt.
libstdc++ errors? Seen them.
Yes! A thousand times yes. I'd rather reinstall every single compromised program than deal with the ridiculous complexity of dynamic linking.
My feeling is that dynamic linking is a pain because sometimes one needs, say, 1.2.5 for some part of the code and 1.3.1 for another part, and they are not compatible. Which hints towards the library author having messed up something, right?
Yay, we get to keep on using old Log4j now because one program holds it back.
Surely, if you care, you have a mitigation plan for this?