Does superpage management get easier when each superpage is composed of only 32x of the next size down? When I first stumbled on this idea, it seemed like it would have many more opportunities for forming intermediate-sized superpages.
* Blow-ups in various kernel data structures. There was some virtio code which was allocating N pages per driver queue.
* Problems with GPUs, either the driver or the firmware assumed 4k pages. (Edit: This actually affected Power, not ARM, but the issue is caused by page size: https://lists.fedoraproject.org/archives/list/devel@lists.fe...)
* Filesystems make assumptions about page size versus block size.
* Processes generally take more RAM, with RAM wasted because of internal fragmentation.
Last I checked, at least CentOS 8 (which should be the same as RHEL 8) is still using 64k pages (search at https://git.centos.org/rpms/kernel/blob/c8/f/SOURCES/kernel-... for CONFIG_ARM64_64K_PAGES=y).
And AFAIK, to access the maximum amount of physical memory in AARCH64 (52-bit physical addresses, instead of 48 bit physical addresses), you must use 64k pages. Since RHEL is normally used on servers, it makes sense to want to be able to access huge amounts of physical memory; that's probably the true reason (or even the sole reason) RHEL uses 64k pages on AARCH64.
You have: 2^48 byte
You want: tebibyte
2^48 byte = 256 tebibyte
2^48 byte = (1 / 0.00390625) tebibyte
how many servers do you have with more than 256 TiB of RAM?2MB vs 4kB isn't quite the same ratio as 4MB -> 32GB, but it's still a lot less pages to cache in the TLB, and it's not too big to manage when you need to copy on write or swap out (or compress with zram) and whatever else needs to be done at the page level.