Linux Kernel Developers Discuss Dropping x32 Support
phoronix.com
phoronix.com
An add takes 1-clock cycle (latency). An L1 cache lookup takes 4-clock cycles (latency). So an add + L1 lookup would be only 5-cycles of latency. Skylake can do two memory lookups per clock and many adds per clock (I think 3 adds per clock?), so it wouldn't take many resources at all.
The actual memory lookup takes ~100+ Cycles if its outside of the cache.
It seems like this problem can be 100% solved in userland without needing kernel support. For something like following a linked-list chain, I'd expect almost no change in speed due to out-of-order executions and memory latency hiding techniques of the CPU.
----------
64-bit user space should remain the default, if only for ASLR / security reasons. Its actually trivial for CPUs to traverse an entire 32-bit space (4-billion isn't a big number: most computers today are ~3GHz, or 3-billion operations per second...)
So if you want your address space randomization to actually have any security what-so-ever, you better use the 64-bit (erm... really 48-bit) address space.
But in this case you can't do it entirely in userspace because things like mmap exist. You get pointers from syscalls all the time, and those can't really be compressed down since there's not a single fixed base address to add onto them. You'd have to use a different type to represent malloc'd memory that is conforming to 32-bit compressed references, but that gets real messy real fast.
Main issue I see is that this requires a new ABI for passing pointers around.
Isn't that what x32 actually is? I mean it's not just for passing pointers around.
Aaaand that's the point. With x32, people who care about these sizes can use the x32 ABI which is consistent in its own way. While everyone else uses the x64 ABI.
Without x32, people who want to have 32-bit pointers need to reintroduce the NEAR/FAR distinction into the source-level-APIs.
Ouch.
I'm guessing the real rationale for getting rid of x32 is that 32-bit pointers are just not a big enough use case to justify it.
However, It doesn't defeat the criticisms that x32 is facing in the linked LKML email: it's implemented in a brittle way; it lacks real users; and toolchains often either lack x32 support or accidentally break it.
A pointer to a data structure is shifted right by 3 a.k.a divided by 8. You now have 35 bits and can addres 16GB. Data structures should be aligned on an 8 byte boundary.
At runtime, the pointer is expanded to 64 bit in a register, multiplied by 8, and the correct offset for a struct field is added.
So you can't point at individual bytes, and you have to shift, but you gain 4 bytes per pointer. It is said this hits the current sweet spot for Java.
Is this still the case though? Does JVM performance fall off a cliff without compressed pointers? Or is the issue more GC beyond 32 Gigs.
Long pauses for full GC are a distinct issue that large heaps have. I believe the newer garbage collectors are designed to address this, but don't have much experience with them in practice.
The heap is limited to roughly 30 GB, plus a bit of overhead. You will never run in 32bits mode if you set a 32GB heap, because overhead.
From experience, a 30GB heap is similar to a 40GB in usable space. 32 vs 64 bits pointers respectively.
https://www.elastic.co/guide/en/elasticsearch/guide/current/...
I understand there's overhead I wasn't suggesting that all 32 Gigs were useable
32 GB, right? 2^32 is 4 GB, so 8 times that is 32 GB, unless I'm missing something?
Edit: yep: https://wiki.openjdk.java.net/display/HotSpot/CompressedOops
More generally I suspect that this approach might make sense for a VM but I that it would lead to huge headaches in C code because the programmer would have to manage the "compressed" pointers explicitly in code unless the compiler was smart enough to handle it completely by itself (something that could be very tricky to do efficiently). It reminds me of the good old days of segmented memory with the mess surrounding near and far pointers.
Still, I would be curious to see actual performance numbers for these various "64bits ISA with 32bit pointers" ABIs, it is true that 64bit pointers are pretty wasteful for many applications.
If you're talking about x32-style mapping then of course that works but then you still only have 32bits of address space.
Having only 32bits of address space wouldn't be a problem, isn't after all the goal of x32 to reduce memory consumption of processes that don't require more than 4GB?
The mapping would just simply take place on a per-process basis.
But that's Java. In C, pointer casting makes it impossible to do this without help from the programmer: You simply can't stop the programmer to cast an invalid value to a pointer. In theory, a mode where char is 2, 4 or 8 physical bytes, but in practice, a lot of software would break.
The main optimisation is that more data fits into the CPU cache so potential for more cache hits which leads to improved speed.
Ideally every application that is unlikely to use more than 2GB of memory should use the X32 ABI.
Cache latency is the most important bottleneck of today's general purpose CPUs. X32 Helps a lot with that and not just in "synthetic" benchmarks.
It is absolutely idiotic to have 64-bit pointers when I compile a program that uses less than 4 gigabytes of RAM. When such pointer values appear inside a struct, they not only waste half the memory, they effectively throw away half of the cache. -Knuth
It's often useful to be able to address in 64bits even if you are never going to use 2GB of memory. And it's probably more and more true the more you care about performance (e.g., mmap() and friends). So I don't think the set of useful workloads that get a performance boost from smaller pointers and don't get hindered by a small address space is as large as you are assuming.
https://link.springer.com/chapter/10.1007%2F11688839_14?LI=t...
found speedups of 13.4% for integer workloads. That's significant for my space, hence my interest.
Using a single 1G page per process means switching is faster, and might have other benefits in a post-spectre world...
I was replying to saying that most apps should use x32 if they don't use much memory. That sounds like a nightmare in a general purpose distribution, with having every library duplicated and interoperability issues.
I get the part about duplicated libraries, but I'm not certain it's that big of a deal: Most people don't have more than 2G of code so it should be possible for the 64-bit libraries to be used (and just mapped low) with an m32-built application, if the things that create new pointers can be limited to 32-bits.
One for x86_64, one for x32. And if you want to support traditional 32 bit application you'll need the i686 variant as well. None of these variants are ABI compatible with eachother. If nothing else, that's a burden on logistics.
The most common use case is to just mmap() a bunch of cache files and just fully outsource the cache hierarchy of those to the Linux VM subsystem which will do a much better job than almost any app even can (because it does global optimization across all processes).
> I get the part about duplicated libraries, but I'm not certain it's that big of a deal: Most people don't have more than 2G of code so it should be possible for the 64-bit libraries to be used (and just mapped low) with an m32-built application, if the things that create new pointers can be limited to 32-bits.
This is way too much complexity. "Things that create new pointers" are any functions that return new memory and there are too many of those. A general purpose distro will never take on that much complexity for a 13% speedup. If they wanted even a fraction of that complexity they'd compile x86_64 at several levels of supported instructions for bigger gains from just tweaking a few compiler flags.
I don't understand what you're saying here.
Why do you think having the pointers be 64-bit instead of 32-bit help if the program isn't using more than 2G?
Can you give me a specific example?
> This is way too much complexity. "Things that create new pointers" are any functions that return new memory and there are too many of those.
mmap() and brk() are the two big ones, and almost everything else derives from them.
The tricky part is to get MAP_32BIT into every call to mmap(). That means hacking up the dynamic linker, or having a personality flag that just pins it kernel-side. I don't think either is onerous though (probably a few hours work).
The program isn't using more than 2G of RAM it just mmap()s more than 2G of files and so needs enough address space to fit those mappings. I've come across that when doing caching for an image processing program. The simple way, that works very well, is to just have the thumbnails on disk and mmap() them in. With even medium sized collections it's easy to lazy load more than 2GB of files but that's fine because 64bit address space is cheap and the Linux VM will take care of throwing away things that haven't been used. Makes for a very simple memory+disk cache hierarchy that the OS manages for you entirely.
> mmap() and brk() are the two big ones, and almost everything else derives from them.
Sure but that's at the base of the call chain. Each library will then bubble up those 64bit pointers to it's callers. You can't have only 32bit versions because at least some part of the apps will need more memory so you're stuck having two versions of everything.
If you need the address space, then use 64-bit.
I don't need the address space.
If that was confusing, my apologies.
> Sure but that's at the base of the call chain. Each library will then bubble up those 64bit pointers to it's callers.
Which will have clear upper 32-bits, which means using %eax instead of %rax in main() will work fine.
Well obviously, but you asked for the use case. I was replying to someone saying x32 should be used everywhere 2G of RAM is enough. And I was pointing out there were other reasons 64bit was required and those can even be embedded in libraries in general usage. So I think x32 makes sense if you have a specific use case where the speedup is relevant but not for general usage.
> Which will have clear upper 32-bits, which means using %eax instead of %rax in main() will work fine.
I don't see how this will work without recompiling every library, which is what x32 does right now I think. Everywhere in the code there will be dereferencing of 64bit pointers and every pointer stored in every struct will be 64bits. So does it really matter if the upper 32bits are empty from a performance standpoint?
Yes!
main() would push %eax instead of %rax, effectively doubling your cache size (in words).
> I don't see how this will work without recompiling every library
The x86_64 abi uses registers for the arguments, so fopen() would receive file in %rdi, and mode in %rsi, but the caller only used %edi and %esi. When fopen() calls malloc() our libc knows we want 32-bit pointers, so it uses mmap(MAP_32BIT) so the pointer has the top bits clear and thus when fopen() returns the pointer in %rax, main() can still store it in a 32-bit pointer reading %eax.
Structs are certainly a problem. If main() tries to look e.g. at struct utsname, it would need to know (in this case) that these are "long" pointers rather than short ones. This pain would need to be handled by main() (for example, using wrappers of some kind), but it wouldn't require building or maintaining two sets of the libraries.
All that complexity is certainly not worth it versus just shipping a second set of x32 libs. At least that doesn't require those tricks and wrappers and is at least easy to maintain within the current multi-arch infrastructure distros already created to support 32bit binaries in 64bit distros.
> you can mmap parts of a big file no problem with large file support in libc.
For the speedups discussed here I definitely don't want to spend any time writing special code just to be able to read a 3GB file. That ship has long sailed and good riddance to all those subtle issues where things seem to work fine until they don't and it's not clear why.
Don't forget to not put the bugs in, though!
And as the kernel is responsible for caching (or what kind of caching do you mean?) You will not lose that.
So all the hassle of maintaining a separate ABI (toolchains, kernel, libraries etc.), for a modest performance increase for some subset of applications?
Starts to smell pretty niche to me.
In principle yes, but it is hard to predict which code will be at some point performance critical and which code will not. In the worst case you chain a lot of alleged "non performance critical" tasks together and in the sum (actually product, performance percentage is multiplicative) your problem becomes a performance problem.
> Or might want to mmap() large files
This is a fringe application that is not covered by my definition of "application that is unlikely to use more than 2GB of memory". Most programs read linear from files.
> So all the hassle of maintaining a separate ABI (toolchains, kernel, libraries etc.), for a modest performance increase for some subset of applications? Starts to smell pretty niche to me.
As far as I see it there is not enough awareness of this feature. So instead of declaring it "niche" that causes "hassles" maybe advertise it more since the hard work (toolchains, kernel, libraries etc.) has already been done. (My personal use case is derivatives market simulations and performance the gain is in the double digits percentage. I'm pretty sure other use cases exists, they just need to found.)
In the end, who defines "hassle", how big is it really, and who defines "niche"? Potentially accelerating 90% + x of code running on your CPU is hardly a niche.
Hardware is typically a price sensitive netbook contrained by RAM and with a low powered x86-64 CPU, so that 13% performance increase might be noticeable.
The software is a design that is inherently multiprocess for security reasons (where every webapp is essentially a browser tab running in its own sandbox).
My guess would be no. For a sort-of counter-example, Apple started migrating to a 64-bit environment for their smartphone and tablet systems already back in 2013, and I think 64-bit Android systems are also pretty common nowadays.
And, if they wanted X32 support, I think it would be more than a "couple of commits in their issue tracker"; AFAIU V8 has no X32 support, which would definitely be a non-trivial undertaking. Do you think they would commit to implementing and maintaining V8 X32 support, just to get a modest performance improvement in such a niche market? Also not to mention that for security reasons they might prefer 64-bit anyways (see ASLR).
My opinion is that 32-bit is rapidly dying for anything approaching "general purpose" (including things like smartphones etc.). For embedded system, sure, 32-bit systems are under no threat of extinction; heck neither are 8 or 16-bit microcontrollers under any particular threat. For an example, look at RISC-V; they haven't even bothered to bring up Linux support for the 32-bit variant, RV64GC is the baseline they are targeting. 32-bit RISC-V seems to be targeting microcontrollers or maybe some simple embedded OS's, but not Linux.
Huh? So if I chain together tasks A->B->C, and assuming (optimistically) that X32 completes a task in 90% of the time of X86-64, then the total time of X32 is not .9x(time_A+time_B+time_C) but, what? .9^3(time_A time_B time_C)? I'm lost, can you explain?
> As far as I see it there is not enough awareness of this feature.
Hmm, I'm not sure. Seems it has been discussed extensively all over the place, certainly enough that people interested and able in shaving some (likely) single-digit percentage time of their application, that they have determined is due to mem/cache usage of pointers, ought to be aware.
> the hard work (toolchains, kernel, libraries etc.) has already been done.
That is true, but the point is that there is an ongoing maintenance burden. If almost nobody is using it, is it worth it to keep paying the maintenance cost? Evidently the Linux kernel maintainers are starting to think that nope, not worth it.
> In the end, who defines "hassle", how big is it really
Presumably the people who have to maintain it?
> and who defines "niche"
Number of users, maybe? Did you read the lkml thread? The X32 Debian maintainer popped in, and mentioned that per popcon X32 in Debian has 7 users. Not 7k, 7. Seven. Now there are certainly users who are not enabling popcon, but still. For comparison, for x86-64 popcon reports 172k users. So for every user who thinks X32 is a good idea and what they need, we have 24k users who prefer x86-64 (certainly a large fraction of those implicitly, but still).
Lets assume Instruction I_n is dependent on I_n-1 is dependent on ... I_1 (I_1->I_2->...->I_n). The Instuction I_i takes t_i amount of time. Lets further assume those Instructions can be pipelined and parallelized hence the number of executions of the full Instruction series is big compared to n. Then the time needed to execute the series is max(t_i). If we have a optimization that modifies all the times to o_i * t_i the new time needed to execute the series is max(o_j * t_j) hence the time ratio saved on the whole series is r = max(o_j * t_j)/max(t_i) and the time saved per execution is then n * r because the time you save to execute the whole series is not dependent on n itself.
> Debian maintainer popped in, and mentioned that per popcon X32 in Debian has 7 users.
In order to do it right you have to distribute X32 versions for packages where it makes sense in your otherwise X86_64 distribution. Then you have increased the performance for everybody and an adoption of 100%.
Seriously, most people I know just don't think that critically about 64-bit, and don't understand the performance gain of pairing 64-bit instructions with 32-bit pointers.
This is a discussion about whether the kernel itself wants to run using x32 mode, which is arguably much less useful. Such a system would be limited to 4G of physical memory, etc... There aren't that many such systems left in the x86 world (that run Linux, anyway -- I'm literally working on an x86_64 port for Zephyr now, and it uses x32).
Such kernel will run x86 binaries (in compat mode), x86_64 binaries and x86_32 ones.
If the kernel removes X32 support, I guess it won't take long for toolchains to remove it either. Why should they pay the maintenance burden of it if nobody uses it?
> This is a discussion about whether the kernel itself wants to run using x32 mode,
No, the kernel has never supported running in X32 mode itself. This is about the compatibility syscalls allowing a X32 userspace with a x86-64 kernel.
So, if some programs in the system use the 64-bit ABI and some use x32, then memory usage may substantially increase. And, of course, this also affects CPU caches.
Does anyone have more details on this?
Notably, this means the actual pointers to that memory are still 64 bit.
Obviously, when you change pointers to be 32 bit, that requires an ABI change because the binary layout of your data structures is different.
You could do everything else in userspace if you really wanted. At the cost of safety, type handling and optimization from the compiler as all it would see is 64bit pointers being typecast and generic 32bit numeric data. Whatever you were looking to gain would be lost in extra work, and you can forget about debugging. But theoretically...
There aren't many syscalls that return pointers; mmap() is one of them, and using MAP_32BIT means that the pointer will be 32-bit.
32-bit code tells the kernel to allocate in the lower 32-bit address space (via MAP_32BIT) precisely so that the 64-bit pointers the kernel works with can be safely truncated to 32-bit user space pointers.
For a start, the kernel sets up virtual memory with stack at the high addresses, which would be out of reach for 32-bit pointers (without translation). Secondly, pointers passed to the kernel would have to be converted to 64-bit and vice versa.
Conceivably, you could remap your virtual address space to have a stack and libraries within the 32-bit space, and then jump into x32-like code which translated pointers to and from 64-bit for syscalls and kernel data structures. This would retain most of the benefits of x32 for CPU-bound processes.
That's what I'm thinking about. You'd need a little cooperation from libc and the dynamic linker in order to have a good "user" experience, and I think you can even share 64-bit libs (provided they don't call mmap() themselves and return the pointer to your application), but these seem almost manageable.
The performance hit can be quite heavy!
Cache is a limited resource, and you want as many pointers as possible to fit there.
32-bit x86 support is for compatibility with older software.
x32 is a completely new thing that has nothing to do with compatibility. Basically, it's everything that AMD added to the 64-bit variant of x86, except the actual 64bit pointers. So you get more registers, access to additional instructions, and so on.
What you are suggesting might work generally if it was somehow worthwhile to try and do some lossless compression of page contents before storing them in the CPU cache. If that's not done yet I'm sure it's because the energy and extra transistors needed for it are not worth it, and not because CPU designers haven't thought about it.
In set associative caches, one or more of the "ways" in each index could be optimised for these relatively common small numbers, with fewer bits given to it, capable of only storing numbers of the small form.
But these aren't caches of 64bit numbers where that optimization would apply. These are caches of 4kb pages of bytes. And since this is x86 I don't think it's even guaranteed that pointers are 64bit aligned. Maybe for the TLB cache an optimization like that would be useful?
This isn't true. x86_64 CPUs use 64-bit cache lines, in associative sets. They don't have any caches which contain entire pages. The TLB stores the mappings for whole pages, if that's what you're thinking about?
Coincidence... I was just working on such an optimisation. But rather than compression as some commenters suggested, the method is to use separate cache lines for the low and high 32-bit halves, re-encoding values so that small signed and unsigned integers and JVM-style compressed pointers (> 4GiB) only need the low half, and keeping track of when the high half is not required.
That saves cache space. Cache bandwidth to main memory can also be reduced, which also reduces energy consumption on the memory bus, but this is considerably more complicated and involves changes to the ECC encoding.
The same method works with 32-bit pointers, non-pointer integers, and JVM-style compressed pointers over a larger than 32-bit range.
These methods don't save RAM, but they do reduce cache pressure and memory bandwidth.
Like it or not, more and more applications are being replaced with web applications, and many of use simply have too many communication channels, and have web apps open for a calendar, slack, email, WhatsApp etc.
>"The x32 ABI allows for making use of the additional registers and other features of x86_64 but with just 32-bit pointers in order to provide faster performance when 64-bit pointers are unnecessary."
Can someone say why would 32 bit pointers provide faster performance then 64 bit on modern hardware? Is this performance difference really non-negligible?
With narrower pointers your memory & caches becomes effectively bigger, as well as memory and cache bandwidth.
The reason why it's faster is because your program is smaller, and your data is smaller, and that means a better (more efficient) use of the on-CPU cache.
The impact of this is AFAIK generally in the single-digit percentage area, but if you have an algorithm that just fits everything into the caches with 32 bit pointers, the change to 64 bit pointers might push it over the edge, which can then result in significant (>= double-digit percentage) performance decrease.
Only if the cache is full of pointers, and not the numbers you're crunching :)
Worst case scenario is probably a GOT or vtable.