Sorry state of dynamic libraries on Linux
macieira.org
macieira.org
Using LD_PRELOAD isn't so rare and strange that you can propose throwing it out without quantifying its performance cost. I can't recall offhand why I needed it (I would guess Valgrind or Massif) but I've used it several times as a developer. What exactly is the payoff for giving it up? It can't be bigger than the current performance difference between statically and dynamically linked executables, can it?
It also looks like PIE+PIC is required if you want a secure system with ASLR: <http://blog.flameeyes.eu/2009/11/02/the-pie-is-not-exactly-a....
Flameeyes (Gentoo dev) has sent patches to all the main library developer that make sure only the necessary number of symbols is exposed and as much data as possible is marked as read-only. I think this effort is more valuable than proposing a very unlikely ABI change.
Edit: rushing out the door here, last sentence was referring to this: http://gcc.gnu.org/wiki/Visibility
Is the address of externalFunction that is stored in the GOT really resolved to the address of the stub and not the final real address of the function? Because it has to be for data (or code simply wouldn't work), and I don't see why function symbols would be any different.
Also C and C++ say that a function has the same address in all translation units, so comparing the address ought to match regardless of PIC or shared libraries or symbol overriding, and if they don't it's probably a bug in the compiler/linker (or you're using a nonstandard option that breaks this guarantee, e.g. symbol hiding.) And actually this should require that the address in the GOT is the final address of the function.
By the way, any CPU that has static destination branch prediction will predict as well for a double indirect call as a single indirect call (assuming of course the addresses don't change.) The cost is taking up two entries in the branch prediction tables, another L1I cacheline for the stub, and a hiccup in instruction decoding which may or may not have a real effect depending on the code before and after.
> If there’s a reason for getting the address indirectly like this, I have yet to find it.
It should be because of PIC, and the fact that PC-relative addressing on x86_64 has only ±2GB displacement, so if your final binary is over 2GB the linker could fail to put the symbol within range of the offset and fail. Whereas for calls the linker can just insert a stub if this happens and noone's the wiser. Disabling PIC results in "movq $externalFunction, externalVariable(%rip)" for me.
But -mcmodel=small is the default, which should contradict this explanation...
EDIT: so I just tried a test and it appears the GOT on Linux really does contain the address of the stub. what the fuck
On OS X it contains the real address.
Until perhaps very recently, the ISO C and C++ standards didn't actually support dynamic linking. (They didn't officially support multithread concurrency either but flexibility is one of those languages' strengths).
if your final binary is over 2GB the linker could fail to put the symbol within range of the offset and fail
I have had to code around this limitation too but it didn't turn out to be that hard in practice.
How about we optimize for the case where the final binary is 2GB or smaller? :-)
How so? They certainly didn't mention it but they shouldn't have to - describing the final linked behaviour is enough and means that whether it was statically or dynamically linked doesn't matter if it produces the same run-time behaviour. Which resulted in a huge mess in the linker for C++.
> How about we optimize for the case where the final binary is 2GB or smaller? :-)
I agree, but compilers should be standards-compliant by default and any such optimizations should be under non-default flags (e.g. -fvisibility-inlines-hidden). But -mcmodel=small is already the default...
Believe it or not, in C++ a pointer to an object or function is valid as a non-type template parameter. Heck, I bet you can even partially specialize on it.
- Huge use of template metaprogramming ?
- Generated code ?
- May be it's only for the debug build with all symbols ?
void func1(void);
void *func2(void) { return func1; }
You'll notice that unlike function calls, the code here differs with PIC enabled. I think your assumption is that function calls and function addresses in C use the same address, which is not true.This is actually the specific thing that prompted the author to investigate and write up this whole post. It's so what the fuck to me that I didn't believe him at all until I tried it myself.
EDIT: actually, I finally realized one potential benefit to this: it lets you completely lazily resolve the functions address. But again, variables can't benefit from this so need to be resolved immediately. I really wonder whether the load-time gain is actually worth the runtime hit for this case...
Google cache: http://webcache.googleusercontent.com/search?q=cache:http://...
G+ discussion with the author: https://plus.google.com/108138837678270193032/posts/No8T7VLo...
A related issue which is more about the compiler than the operating system is that PDBs with Visual C++ on Windows are much more useful with highly optimized binaries than symbols (either embedded or separate) are with GCC-optimized binaries. This is understandable when you consider the different development cultures, but that doesn't make it any less of a problem for us. :)
99.9% of the time we don't need to, but in that 0.1% it makes a big difference.
Then you can open the core file as usual: gdb executable core (just be sure to have debug files in the same directory)
you just strip the debug symbols out (and put them somewhere safe). then write a .gnu_debuglink section to the stripped ELF binary with a CRC that matches the stripped symbols.
once something bad happens: you just take the core dump, the symbols you have tucked away, and you are able to debug just fine.
http://code.google.com/p/google-breakpad/wiki/LinuxStarterGu...
I think that things like RPATH and LD_PRELOAD are exactly what shared libraries should be doing. The reason for shared libraries in the modern age, is increased flexibility.