ARM32 Page Tables
people.kernel.org
people.kernel.org
This would reduce the amount of bookkeeping if the OS hands out physical memory in large contiguous tracts, but increase it if it is shuffling around individual pages. I have no idea what OS behaviour here actually is.
> This would reduce the amount of bookkeeping if the OS hands out physical memory in large contiguous tracts, but increase it if it is shuffling around individual pages. I have no idea what OS behaviour here actually is.
On x86, large 2MB and 1GB __cannot__ be paged. The latency associated with a 2MB or 1GB read/write to disk is just too long to be reasonable.
These are called "huge pages" (2MB or larger). Memory-intensive applications see benefits from huge-pages, but the additional book-keeping to the user (in particular: huge pages CANNOT be paged, and therefore are a "rare-ish resource") has made HugePages somewhat uncommon on both Linux and Windows. (Windows only supports the 2MB hugepage, and ignores the 1GB option entirely)
> This would reduce the amount of bookkeeping if the OS hands out physical memory in large contiguous tracts
Note that all modern chips have a TLB-cache, to the point where the bookkeeping is largely handled by the CPU itself. This TLB-cache still has restrictions, so a huge-page reduces book-keeping and accelerates the speed of memory access slightly.
The OS sets up the table, but the CPU itself handles the whole shebang nearly transparently. That's why all of these memory-address and data-structure specific details show up when talking about page-tables, you've got to set up the data-structure precisely so that the dedicated hradware inside the CPU can traverse the table correctly.
On spinning rust, seek times are O(10ms) and speeds are O(100MB/s). That means in the time it takes to seek to page in a page, you could read about 1MB worth of data. That should make that order of page size vastly more efficient for paging in data than tiny 4kB pages, from an HDD.
From an SSD, sure, the seek penalty is much lower so things should skew in favor of smaller reads.
With server class spinning disks, you get something like 90% of sequential bandwidth when doing 8 MiB random reads. There's slight gains up to 32 MiB. So yes, overall bandwidth out of the disk is much better with 2 MiB than 4 KiB.
But the question is: how much of that 2 MiB block is actually useful before it gets ejected out of the page cache? For 1 GiB of page cache memory, with 4 KiB pages we get ~260k different entries, but with 2 MiB we only get 512. So we'll end up with a lot of capacity conflicts on the cache unless our workload just happens to only work with a small number of huge memory blocks. That's likely to dominate performance even if we got "free" bandwidth with the large pages.
All that said, 4 KiB is a bit painful on modern hardware. Current SSDs are happiest when writes are an integer multiple of around 32 to 64 KiB or so. It'd be nice if we had a medium size pages option.
On iOS and macOS on Apple Silicon, the used page size is 16KB across the board.
On RHEL/CentOS on Arm64, the used page size is 64KB.
In other cases, it's generally 4KB (and is always 4KB on Android or Windows/arm64).
Odd that I can't find anything instantly. - ok, they're called segments it seems.
Look for the word "segment" here, you'll see it immediately. http://homepage.divms.uiowa.edu/~jones/opsys/notes/22.shtml
Even without what marcan_42 said, this is still false: a 2MB page is made of 512 4KB pages; page one of those out and use the freed memory as a page table for the other 511 (then page those out individually). If you have a contiguous, aligned 2MB block, defragment physical memory and fuse it into a huge page. It may very well be the case that linux chooses not to support this, and that's plausibly the correct decision, but it's not remotely impossible.
Do you have a link where I can read about this? Couldn't find it quickly.
Windows has Large-Page documentation: https://docs.microsoft.com/en-us/windows/win32/memory/large-...
Linux kernel also has docs: https://www.kernel.org/doc/Documentation/vm/hugetlbpage.txt
In my experience, most Linux distros have huge-pages enabled, but just doesn't use them unless the user sets some configuration flags in various locations.
More recent revisions of the architecture (POWER9) have shifted towards a more typical radix page table.
The book Operating Systems: Three Easy Pieces has a section to answer this very question.
I don't think fragmentation is the issue.
The OS doesn't search the page tables. The __CPU__ does, EVERY single time any program does a memory lookup (load/store operation), the CPU may have to traverse the page table. Page-table entries are often cached in the TLB-cache, but if the cache is full, a full table-traversal is necessary.
Its clear that page-tables are optimized for minimum latency: so that the smallest area of the CPU can be used, and the simplest hardware can implement the page-lookup procedure.
"blah = blah->next" is a very common pattern in programming. Since the page-table implementation in the CPU can speed this operation up (or slow it down), its extremely important to optimize for absolute fastest latency.
-----
A lot of stuff goes on with "blah = blah->next". If blah->next is outside the TLB cache, the CPU will have to jump to the page-table, traverse the 1st, 2nd, and 3rd levels of paging (looking for the physical location of blah->next).
A "variable length page" would complicate this lookup severely: the CPU wouldn't know which offset to lookup pages, and your non-TLB pages will be much much slower as a result.
x86 "huge pages" are based on only searching one (1GB) or two (2MB) layers of the page-lookup procedure, instead of all three (going down to 4kB pages). That's how you get support for other sizes without variable-length entries in the page tables / page directory table / etc. etc.
Edit: you would also then suffer memory fragmentation, not just lookup latency.
For already mapped pages latency would be an issue as you described.
[1]: https://developer.arm.com/documentation/ddi0333/h/memory-man... This is specifically for the old ARM1176, but there are only minor page table format tweaks in ARMv7-A and ARMv8-A AArch32 short descriptor page tables. ARMv7+ also supports what they call Long Descriptor Page Tables, which are fundamentally similar, but have 8-byte page table entries, support larger physical addresses, and have a different set of available page sizes, and lots of other little details are different.
Also, why count the directories from the bottom instead of the top?
You count directories from the bottom because you may need to add another level in future hardware to allow handling greater maximum virtual addresses. Choosing a naming convention where extensions increase in numbering is generally a better idea otherwise you end up with a travesty like x86 privilege level (ring) numbering which used lower numbers (ring 0) for higher privilege, so when they needed a higher privilege level they needed to add ring -1, -2, -3, etc. This also is in contrast to how ARM uses higher numbers for higher privilege (EL{x}), so could add EL2, EL3, etc. when they needed higher levels.