The trouble with 64-bit DMA in Linux
lwn.net
lwn.net
Also for everyone else - just a reminder, LWN is a great resource and is funded through subscriptions. Personally, I've found the information on this site to be extremely helpful to understanding how the kernel works - to the point that multiple products I've built were only really doable because of LWN articles explaining the subsystem I'm working with (namespaces, io_uring, etc are all are the topic of fantastic articles there). A personal subscription is extremely affordable and very much worth it. They have corporate site license subscriptions too.
The release of the 68020 and Macs with more RAM resulted in the need for a whole transition to “32-bit clean” system and apps that took a while. I also seem to recall something about early versions of Microsoft Mac apps being limited to the lower megabyte of RAM for a similar reason.
ARM calls it Top Byte Ignore, it's supported on all 64-bit ARM CPUs, and it's been enabled on Linux for a long time (see https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin... and https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...).
Intel calls it Linear Address Masking and AMD calls it Upper Address Ignore, however these will be available only on very recent hardware, and the Linux patches for them are still under development (see https://lwn.net/Articles/902094/ and https://lwn.net/Articles/888914/).
The primary example is Lua (before 5.4, and also LuaJIT): they use NaN-boxing to make their object take only 8 bytes. Although a double (Lua's only number type) already takes up 8 bytes of memory, there are some bits residing on the high 16-bits that are free to use when the number is NaN. With that you can store the type tag in there, which enables you to use the remaining bits for things other than double (like for example, storing memory addresses...) The Crafting Interpreters book has a much better explanation of this: https://craftinginterpreters.com/optimization.html#nan-boxin...
You can look up at https://en.wikipedia.org/wiki/Tagged_pointer to see some more examples. Apple seems to use this a lot for their Objective-C runtime on iOS where the leftover bits are used for things like refcounts or various boolean flags for their objects. Personally I had experience exploiting this in my game engine: for storing the generational index for a memory arena (to catch any use-after-free errors).
[1] https://www.folklore.org/StoryView.py?project=Macintosh&stor...
If only a few issues were noticed, those device dricers could likely be fixed with small code changes and most people would be happy. When Linus's machine breaks, there's a heuristic that many machines will break, so that's a recipie for a release that will cause lots of upset.
HPC has been moving huge data cross PCIe running Linux for a long while, there must be multiple ways to overcome IOVA-4GB-allocation limit.
Last, they can be dealt with at hardware level, let the (smart) hardware(e.g. an offload engine, or accelerator) to manage the memory on its own for DMAs.