32-bit x86 Position Independent Code – It's That Bad
ewontfix.com
ewontfix.com
It looks really terrible to lose one precious register and have all this PIC overhead.
I don't use mono's ahead of time compilation because it also creates PIC. The JIT has one register more available. But I haven't measured it yet
ELF actually supports (and always did) non-PIC code on x86 Linux in shared libs, ld.so will do fixups to the code on the fly. See http://eli.thegreenplace.net/2011/11/03/position-independent...
Among other things, because that means the text section has to be writeable, and can't have a shared mapping across processes unless it's mapped in the same place in every process, which breaks address-space layout randomization (ASLR).
If multiple processes load the same DLL and the default location happens to be chosen well and there is virtual address space in every process, then the linker can simply share the DLL.
That's how I think it was before ASLR came around.
I think Windows removes the writeable-flag again after relocating.
For the area I described (Mono's JIT vs. Mono's ahead of time compilation as PIC .so-file) I personally don't care that much for ASLR, because it's a managed language and because of specific circumstances of that project.
The whole process image is writable when the loader is running - else how could it load the code into memory? :-) Section attributes are applied after loading.
(or -R for short)
In fact, if you call LoadLibrary at runtime, there is no symbol resolution; you have to look up symbols yourself using GetProcAddress.
Under dlopen, symbol resolution takes place; it can change the destinations of function calls.
I'm probably a bit biased since I'm rather used to Windows' explicit import-export system, but having function calls change just because a library was loaded sounds like an opportunity for some extremely confusing bugs.
Funny how perceptions change depending on the angle you're looking at a problem.
Dynamic linking makes security more difficult to reason about, which is a Bad Thing. It has many pluses, but also many minuses. And with symbol versioning it gets very problematic to actually figure out what your real code path is to begin with.
And you wouldn't need to restart the clients with static OpenSSL — you need to recompile them. And with static linking, I'm not sure how you would easily determine the linked version out to say, the minor or the micro. (Perhaps, if this is your OS's thing like nix, the package manager keeps track…)
http://manpages.debian.org/cgi-bin/man.cgi?query=checkrestar...
https://gehrcke.de/2014/06/good-to-know-checkrestart-from-de...
It not only shows processes that run older solibs, but in the case of services, it will also give you the commands to restart them :).
You only need to relink the apps, unless the newly patched library breaks its own ABI. Otherwise, it would be ideal for package distributors and package management systems to ship just a single .o file for an app, and do the final linking at package install time. Then updating a buggy static library doesn't require full recompile of any apps.
Nix generally uses dynamic linking, though the links are to absolute paths to a specific library version, so upgrading OpenSSL requires a recompile much like with static linking.
With "normal" dynamic linking, upgrading OpenSSL and restarting some service means you're now running untested code, hoping that the new version of the dynamic lib really doesn't change an interface on which your code depends. But to be fair, this is usually a good assumption.
There is no need for dynamic linking for updates. Really.
Also - weakly related - what if you want to use a plugin from vendor A inside the program of vendor B, and they live in the same address space - how are you doing it without dynamic loading?
The proprietary code isn't going to chance using system libs beyond libc, because of portability. So they likely still need to be patched separately.
Then don't do that. Seriously, RMS, ESR and others have written a plethora of well-reasoned essays indicating why proprietary software is a poor choice.
Well, I'm convinced.
'Well, don't stab yourself in the eye.'
Even if you so far as control the instruction pointer as an attacker, you might just not know where to jump to in the target address space...
Of course being an ABI issue, probably makes it really hard to solve, and hence hacky solutions abound. (actually its probably pretty easy to solve given LLVM, just impossible to get anyone to accept).
So yes, you'd have to create a new runtime from the ground up, not just a new ABI for an existing API.
Yes, it is a mitigation strategy that wouldn't exist in an ideal world. But we have to do best with what we have, and security takes a very pragmatic approach at this.
For every package, you link statically, you as the parent package owner become responsible for all security flaws of all the packages you link statically.
And by "responsible" I mean: Every dependent security announcement of a dependent package also becomes your security announcement.
If you're AWESOMY 1.1 and you link statically against openssl and openssl announces a security flaw, then you better and quickly release AWESOMY 1.1.1 with an accompanying security announcement too.
Are you willing to do this? Do you trust the chain of dependencies all the way down to also be willing to do the same?
As a responsible developer, I'd much rather delegate that responsibility away to a packager or even the user, especially with well-known libraries like openssl.
If I'm owning AWESOMY 1.1 and link dynamically against openssl, then I don't have to do anything when openssl releases a security announcement. I can, if I want to, inform my users to maybe update openssl, but with some likelihood they are already doing this anyways for some other package.
For me as a developer, this is considerably more convenient.
Yes. Static linking has huge advantages for me as a developer too, but it also comes with a great many additional responsibilities I'm personally not willing to take on, also because I don't trust my dependencies to be as diligent about their dependencies.
Dynamically linked, dlopen()-or-equivalent plugin designs are an exception. Otherwise, I'd like to see more statically-linked applications.
Well, speaking as an administrator rather than a developer: a developer shouldn't be making that kind of policy decision (time of symbol resolution) to begin with barring a strong technical need (plug-in based architecture, etc.).
I mean, 90% of packages don't care whether their libraries are dynamic or static, but the ones that do can still be annoying. (Coreutils, of all things, requires dynamic linkage for stdbuf, though that seems to be considered a bug by the maintainer; on the other end of the spectrum, getting Perl to even build statically is like pulling teeth, and that was a design decision.)
Do you trust the chain of dependencies all the way down to also be willing to do the same?
The only time that is a problem is those cases where the package developer literally plunks down a bunch of code from his upstream into his own tree (think the embedded glib inside pkg-config, or the hideous monstrosity that is gnulib). I think all sane people agree this is bad -- if you aren't significantly altering the code (at which point it's "yours") there's no sense in doing that and risking a sync problem. But that's not even static linkage; that's just literal code-sharing.
That said, I'd like to point out that it is convenient for the developer - since customers are ultimately interested to know if your software is vulnerable, and you'll have to explain how the vulnerability affects it -.
It's also a double edged sword: openssl has a good ascending compatibility record, but that can't be said about all user space libraries. IOW, it's tricky to guarantee that your software will work flawlessly across all the incarnations of CentOS 6.x, if you have a lot of external dependencies, for example.
It's not so tricky, since Red Hat specifies nowadays what guarantees can be expected for what packages:
https://access.redhat.com/articles/rhel-abi-compatibility#Ap...
I have absolutely no need for my authentication system to be "flexible".
Yes, and no, because as the developer you can look at the security vulnerability and decide if its actually exploitable in your application. That is assuming you can determine it, but in a lot of cases its simple. Especially in a huge library like openSSL. Say for example the only thing i'm using openSSL for is some limited functionality, say SHA256, then I probably can ignore 99.99% of the security issues because they just won't apply.
I've been in this situation with an embedded platform that ships as part of the product I work on. Its pretty much got daily security updates, and yet we rarely get hit by any of them because our usage of the platform is like 1% of its functionality.
To calculate exactly how much memory is saved by shared libraries, you'd need to write a kernel module to walk the internal structures describing physical pages and summing reference counts of used pages. Maybe it's already been done?
Depends. If I have 7 instances of my terminal emulator loaded (say I hadn't discovered tmux yet or something), the loader can share their .rodata and .text segments, and in many cases it does (YMMV; heuristics apply; void where prohibited; etc.). So a lot of things that are in shared libraries right now might still be only loaded into memory once if their binaries are segmented correctly.
The Plan9 people (Plan9 doesn't do dynamic linking) claim that the memory savings they get from skipping the relocation overhead are greater than the memory hit from the times that the same stuff does get loaded multiple times, though obviously always take self-promotion with a grain of salt.
Use case probably matters -- my servers run few processes to begin with, and it's often a lot of versions of the same process, whereas my laptop runs a ton of very different processes. Same sort of argument that makes me happy with udev on my laptop while also very happy with a static /dev tree on my servers.
[0] In fact, certainly. Program loading works by mmaping the executable with MAP_SHARED flag into the process's address space, and the VFS takes care to keep all mappings coherent. The simplest way of keeping mappings coherent is to actually share the backing storage between all mappings. It's the foundation of CoW.
Unless there's a tool which makes this extremely simple (maybe Intel's pin?), I believe that writing a kernel module is simpler. The module's init function tallies up the pages and writes out the result into the kernel log. Then the module exits.
The results of mincore() in one process with a shared mapping of a file are enough to tell you how much of that file is loaded shared system-wide.
That said, QT and GTK probably should be broken into smaller libraries in a perfect world.
I'm joking, of course, but there IS a lot of variation in computer hardware, which makes this kind of broad generalization even more problematic.
Also, memory size is probably not that useful of a metric for modern hardware, where the penalty for overflowing the CPU caches can be huge.
edit:
Why statically link when you can prelink(8) instead?
This is specific to i386, as far as I'm aware. And even then, the worst parts of it are only specific to i386 using the same SysV calling conventions -- ARM has PC relative addressing, as do most other processors commonly used in embedded systems. If you have PC-relative addressing, the entire problem goes away.
Coincidentally, Android defaults to PIC in newer SDK versions. So there is still a pressure to optimize for IA32+PIC.
http://www.ee.columbia.edu/~mgseok/pdfs/phoenix_isscc_dac_de...
I think the Solaris linker took some interesting approach with shared libraries at runtime as well.
Of course, the designs and implementations have not. They are still in heavy use, 20 years later. What "just works" is never questioned. I am glad to see some people questioning the old assumptions.
Unrelated question: Why does Linux pass arguments in registers instead of using the stack?
It's a convention. Depending on your architecture, compiler and API, it uses the stack in a lot of cases (cdecl/stdcall on i386). Linux on AMD64 always uses the SYSV ABI which uses registers as much as possible (i'm assuming for performance reasons, since 'fastcall' does the same on i386).
http://en.wikipedia.org/wiki/X86_calling_conventions#System_...
Does Minix use that convention? How about BSD?
http://en.wikipedia.org/wiki/X86_calling_conventions#System_...
Anyway, the original question was _why_ does Linux use registers to pass arguments. Arguments can be passed on the stack, but Linux chose to use registers. Why this choice over the other option?
Do you know the answer?
Just curious.
If I were a sysadmin, and maintaining the computer were my job, this would not be a problem, because figuring out why things don't work and fixing them is what I would be trying to use the computer for; but as an end-user, I just want things to work.
If every application were built with static linking, and system libraries only updated with an explicit system upgrade, I would be much more likely to upgrade frequently, because it would be possible to know and limit the scope of churn implied by any given upgrade act.
Others seem to be suggesting that the emergence of dynamic linking was a response to general limitations in secondary storage and memory. Limitations that no longer exist.
I like marssaxman's comment.
I'm not sure I would call this a "feature" because I think of "features" as being intentional and if marssaxman is correct, this justification for shared library use was an accidental side effect.
But I have limited knowledge of the history behind dynamic linking. Hence my questions. I am just curious.
If the functions tend to perform operations that are dependent on the arguments first (I know lots of functions I see do, doing things like an immediate null pointer check and pointer dereference on an arg) then it is better to already have them in a register, you can often immediately replace the value of the pointer in the register with the value from dereferencing.
For your point about there being more complex functions than simple ones, it doesn't matter which there are more of, it matters which are called more often. If every complex function on average calls 0.5 complex functions and 5 simple functions, you still probably have more simple function calls overall.
It's true that the benefit isn't enormous.
Shouldn't that be "foo: call bar"?