The Arm64 memory tagging extension in Linux
lwn.net
lwn.net
By the time the Mac II came out (68020 with 32 bit address bus), it was apparent that this was going to be an issue, and there was an effort to remove this logic. Those early machines were a combination of ROM and loaded OSes, so there was an effort to clean the code for both the distributed OS and the stuff included in the ROMs of the machines.
I think if memory serves the Mac IIcx from around '89 or so was the first machine with 32 bit clean roms, and with the right OS version you were good for >16mb of RAM.
I've also got a hazy memory of some architecture (i'm going to guess DEC Alpha) which only read from 32 bit boundaries, and where the bottom 2 bits of the address were masked out, so you could store data there too...
Anyhow, the real take home from all that was that misusing addresses to store additional information, which might seem like a clever trick is likely to bite you at some point in the future.
Plenty of dynamically typed languages have used the bottom bits in an address to store tags. You have two available on a 32 bit machine, three on 64 bit. The SPARCv7 has instructions that will do arithmetic on tagged values.
IIRC Excel went a bit further by using the leftover bits and even bit(s) starting at bit 24. That worked until hardware stopped wrapping around memory. Apple created a launch error “program had special memory requirements” which actually meant “this is Excel version so-and-so”.
And yes, that can bite you in the future, but the gains today can (appear to) outweigh that.
It also is perfectly valid when you control both the side that determines the constraints and the side that runs extremely close to them. For example, Cocoa’s tagged pointers (https://www.mikeash.com/pyblog/friday-qa-2012-07-27-lets-bui...) are fine, given that Apple both controlled the memory allocator and the library assuming it aligned allocations (they also are safer in that they won’t try to read from addresses that don’t point to valid memory. That’s essential for their use)
But also this is a hardware feature so the code will continue to be compatible anyway.
Even if this acts like a permanent truncation of the address space by 4 bits, so what? If you think you can keep Moore's law going strongly enough to hit that cap, you're probably going to hit the 64 bit wall soon enough after; might as well focus on that shift to 128 bit addresses
https://www.usenix.org/system/files/login/articles/login_sum...
It's faster then doing the same in code (e.g. via 'tagged handles': https://floooh.github.io/2018/06/17/handles-vs-pointers.html), or memory debugging tools like Valgrind or clang ASAN (and unlike Valgrind or ASAN, the checks are always enabled, also in the "release version").
1. It still helps with defence in depth
2. The threat of detection prevents attackers that need 100% stealth
3. Failures are likely to be detected which is still valuable as you can now mitigate damage
4. Failures are likely to be detected which helps get software faults fixed - which allows the ecology of security to improve over time
Anyone needs more than 256 Terabytes of RAM?
In any case, I'd be quite surprised to see a 64-bit CPU doing such heavy lifting in 2040.
I can't imagine anyone needing more than 4PB in the next 30 years, and the security gains now is (in my opinion) worth the potential refactoring in 30+ years that would be required.
Even so, it wouldn't be that hard to actually build a system that could in theory try and map 4PB of storage via mmap today, its about 1/2 a rack of fairly common equipment. Given there are various companies selling 50T ssd (https://www.anandtech.com/show/11639/viking-ships-uhcsilo-ss...) its probably possible to do it in under 10U.
Intel added support for 5-level page tables (bumping virtual memory to 2^57 bytes) a few years ago: https://en.wikipedia.org/wiki/Intel_5-level_paging.
EDIT: My math is bad, see comments below.
I can unsarcastically say that no one will ever need full 64bit addresses on this planet.
2^71 bits - total hard drive capacity shipped in 2016 [0]
I doubt your calculations.
[0] https://en.wikipedia.org/wiki/Orders_of_magnitude_(data)
Google plans to eventually make MTE a compulsory feature.
https://source.android.com/devices/tech/debug/tagged-pointer...
Eg. Store pointers to all these objects into a std::set? The memory tag will now form part of the keys to this set, and would affect iteration order etc... It could cause breakage in existing applications.
There are also cases of custom suballocators or arrays of objects - Looking at an address makes it possibly to figure out which array it belongs to. This code would break.
Granted, it would still be possible to do all this if you just mask off the tag bits, but it requires a software change.
It's always the case that some software that does things that are not valid-by-the-language-standard might break if run on a newer version of the OS or a newer system library version (remember the big flap about glibc memcpy() changing its behaviour when called for overlapping regions?). You don't want to break lots of software gratuitously, but sometimes the tradeoff is worth making.
I'm pretty sure this memory extension doesn't affect uniqueness of pointers... that would indeed break a lot of software ;-)
What I mean is for example: you get a pointer by allocating memory, then free that memory, then you allocate again and get a pointer to the same memory location. Without a tag those two pointers would be equal, even though they come from separate memory allocations. This can be a source of bugs if the pointer is also used as some sort of object id, and I know that at least I stumbled over this embarrassing problem more than once in C++ until I switched to tagged handles :)
Exactly, but just declaring that a pointer is "invalid" doesn't help much with debugging such dangling pointer problems, while tagged pointers do.
It's not like an "invalid pointer" is any different from a "valid pointer" when you look at it. With the extra tag bits, an invalid pointer can actually be identified as such.
Std::set uses it of course.
I did propose storing type values in a TLB a long time ago but haven't run any simulations for it. Had started building some hardware to do type checking on an Atari ST but didn't finish it.
Using multiple RAM chips to increase throughput is an old computer engineering technique; it was, for instance, the entire justification for planar graphics modes in early graphics hardware.
https://www.kernel.org/doc/html/latest/core-api/protection-k...
iirc it was used in lisp compilers/interpreters to speed up object unboxing (as in answering the question: "what class does the object pointed by this pointer belong to?")
Of course, this doesn’t stop you from being able to[a], it just makes it harder.
[a]: The processor doesn’t know the “type” of a register’s value; It’s just bits. It could be a pointer, integer, floating point, etc. until you tell the processor to do something with it. So one could obviously use those extra bits to store data, then when a dereference is needed, store those bits in another register, dereference, then put the bits back. But that’s a lot of work to save a byte or two.
[0]: Intel SDM Volume 1 §3.3.7.1:
> In 64-bit mode, an address is considered to be in canonical form if address bits 63 through to the most-significant implemented bit by the microarchitecture are set to either all ones or all zeros.
Is it? It's something like two xors on a register. If that's a byte or two per pointer (ie: it's pointer metadata that you'd have to store anyway, which I assume is the use case) that's actually a pretty significant win, and could help reduce cache pressure, thereby buying you back the performance (and then some).
Addresses are also XOR'd all the time with a secret as a mitigation, I don't know that there's much of an impact at all.
Emacs Lisp is an example; here's its tagging apparatus.
https://git.savannah.gnu.org/cgit/emacs.git/tree/src/lisp.h#...
Its one of those lessons that everyone seems to be constantly relearning. Address spaces grow.
One of my first "real" jobs, when this topic came up a very senior person said something to the effect. We continue to find ways to use up one of those bits roughly every year. Moving from 32-64 bit VA's is going to get beyond the end of my career but someone will have to deal with it in the future. So its a little slower than 1 bit a year, but everyone is adding a few more bits (52/56/etc) because there are actual commercial computer systems with a few tens of TB of ram today. If persistent memory ever becomes a thing, someone _WILL_ want to mmap their multi PB storage cluster, and then there will be a scramble for more address bits.
You may be thinking of the fact that a full 64 bit VA space isn't supported in current implementations. But this has no impact on use of tagging compared to most other archs.
MTE, OTOH, was specifically designed for detecting temporal bugs. For example, a freed object has the allocation tag (stored in memory) changed by the heap allocator so that the original pointer (with the original tag) can no longer access it. Of course, the trade-off is the 4 bits per 16-byte granule that need to be stored somewhere and the probabilistic nature (1 in 16 chance of hitting it).
But I think CHERI and MTE would complement each other nicely.