https://asahilinux.org/2021/03/progress-report-january-febru...
Userspace breaks on 16K pages when it tries to do things like call mmap() with virtual addresses that aren't aligned to 16K. Usually it's allocators doing this when they think the entire world is 4K.
Don't 16K pages "exist alongside" 4k pages too, just not within the same virtual address space or (in Linux) VMA? How else are Rosetta apps supposed to work in Mac OS?
Rosetta runs the userspace half in 4K mode, and XNU had to be reworked a lot to support this. Linux could of course be reworked to do something similar on paper, but it's a hugely intrusive change and it'd actually be easier to just make the kernel support 4K/16K pages in a single build first.
Hugepages aren't like that, they actually coexist with normal pages. In general, hugepages are just a pile of contiguous/aligned small pages that the kernel manages as a unit, and it flags them to tell the MMU "I promise these are all one big contiguous chunk so you can optimize it to one larger TLB entry". Depending on the page table structure they might be coalesced to a higher-level page table entry, skipping a page table walk level.
This looks like the real issue, so adding support for both "4k" and "16K" address spaces would involve support for multiple page table structures within a single kernel? Still seems very much worth doing since it can likely be extended to support e.g. 64K. And maybe other architectures could reuse that support depending on how their hardware support for multiple page sizes works, e.g. https://en.wikipedia.org/wiki/Page_(computer_memory)#Multipl...
Lots of things in the kernel count sizes in pages. If your page size can vary, suddenly a lot of kernel constants become boot-time variables. And if it can vary from process to process, suddenly lots of things are per-process. Say you run a 4K process. It wants to map some data from a file. That data is in the page cache in 16K chunks. Now you have one page cache page mapped to anywhere from 1 to 4 4K pages. How do you keep track of that? That wasn't necessary before.
What happens if a 16K process shares memory with a 4K process? If the 4K process sends the 16K process a 4K page, that page can't be mapped at all.
See how this is makes everything much more complicated?
Has this stuff been discussed elsewhere so far, e.g. on some linux kernel dev list? I think you've made a good case for not trying to support per-process page size right away, but many of these issues are not entirely new; they came up in some form as part of the transparent-huge-pages feature. It turns out that some hardware support already requires the kernel to understand "higher-order" mappings of contiguous physical pages, and "transparent huge pages" could leverage that support.
In an earlier post they said this about progress on the iommu hw limitation leading to start out with 16k page size: "Sven took on the challenge and now has a patch series that makes Linux’s IOMMU support layer play nicely with hardware that has an IOMMU page size larger than the kernel page size" (https://asahilinux.org/2021/10/progress-report-september-202...)
(Also apps can always opt-in to bigger pages using the existing mechanisms)
Of course it's also true that 4k is a ridiculously small page size, we've been using the same size for 30-40 years while memory sizes have grown 5-6 decimal orders of magnitude. But from that POV we should now do a bigger bump than 4k->16k - I think the next bigger commonly used page size on ARM Linux has been 64k.