Historically ASLR on Apple’s platforms has had many weaknesses, mostly stemming from poor randomization across many different subsystems :P
What I’m also really curious to know is what the performance implications will be of isa PAC…
Historically ASLR on Apple’s platforms has had many weaknesses, mostly stemming from poor randomization across many different subsystems :P
What I’m also really curious to know is what the performance implications will be of isa PAC…
Safe languages don't magically make all code impossible to exploit, they just reduce the attack surface to logical errors, instead of having to deal with UB and memory corruption as well.
The same way that helmets and belts don't save people from dying all in all kinds of accidents, yet they surely help reducing the mortality numbers.
I enjoy languages with helmets and belts.
There isn't a "popular" CPU architecture that is 100% PIC for all of the allowed address ranges. There are "optimizations" the compiler can choose for limited addresses "reaches". Please consider independently compiled libraries that are aggregated into something larger, specifically languages that permit subclassing and inheritance. How can they be moved within an address space?
If there is a hierarchy of address availability this means that ASLR cannot be implemented in a pure way - it must involve runtime fix-ups if the base address changes. If the length of an instruction must change because of the addressing mode, what happens? How can you do ALSR? Do you rewrite all instructions (is there space? was space allocated?) or always use the most expensive address mode?
Languages that support subclassing and inheritance don’t represent any special hazard to relocation; ultimately these are built out of data and function symbols that need to be supported even if you’re only compiling C.
Randomization is NOT free. Randomization either requires that the values are pre-computed before install (which means delivering and pre-computing N different versions) or it is computed it on device.
If the randomization is computed on-device how to you validate that the binary or a library has not been "substituted" - persistent malware, APT?
The "compute on device" was a feature of very old macOS versions - it was annoying and took quite a lot of resources.
"TOTAL ASLR" depends on a CPU arch if it fully endorses over all addresses position independent code and data (Q: homework for ARM, x86_64...). If the ABI allows violations of this you cannot glide / slide all code and data addresses without significant runtime costs. This will likely result in a compromise.
To be clear, I am not complaining about a lack of ASLR where it would be prohibitive, such as mapping the shared cache at a different address for every process (which, unless done carefully, would kill the benefits of it being in shared memory as the pages would all be dirtied). I am talking more about various instances where Apple has generally used very poor slides for reasons that aren't all that great, leading to the randomization being easy to break.
Assume foo.dylib and bar.dylib are system libraries, both live in the shared cache, and foo links to bar. Both are loaded and mapped to a running user-space process.
If foo links to bar, then there must be a symbol table somewhere in physical memory with an entry that points to bar’s TEXT.
That symbol table is part of the shared cache, right? Doesn’t it already follow that bar’s TEXT needs to be at the same virtual address in every process?
Yes, and this is how the shared cache works. If you wish to map the shared cache elsewhere there would have to be another copy of it in memory, which is why this would be a massive pessimization if done without designing for it. Perhaps you might have some idea as to what would need to change to make this not be as bad.
OpenBSD relinks the kernel with new randomization on every boot ("KARL").
Fun stuff, but probably a nightmare to make it work with signature checking :D
My comments were entirely about the shared cache region and how it can be moved.
I again ask you to do the homework - please calculate how it could be done better given the address spaces involved. Address spaces going from 32bit to 64bit have gotten better, but this does need to be kept into consideration given the size of the object involved and the API / CPU instructions available (please, I ask you to consider all of the addressing modes available to the ABI for all of the currently supported platforms)
[added]
It is likely the individual objects that compose the share cache region are compiled independently. There are lots of individual objects! Resolving the dependencies of the composite shared object are likely expensive.
Historically I observe that the contextual data available to a static or dynamic linker has been very constrained, which makes relinking / reallocation objects a challenge.
Oh, and because you mentioned it: yes, the shared cache has many objects in it. The process of making it is fairly involved, especially since many optimizations go into it (string deduplication, perfect hashing for Objective-C runtime metadata, shortening intra-cache procedure calls, …) But this is all done once when it is built by Apple's B&I, so it's not a problem on-device.
There has to be some reason.
And, I am sure there is a reason, I just have not seen anything that indicates anything good enough that it is worth drastically reducing the quality of the ASLR they provide.