How Bad Can 1 GB Pages Be?
pvk.ca
pvk.ca
If either of those two assumptions are false, then the TLB miss times are swamped by the page in time. Paging a 1GB page in is _NOT_ a fast operation, especially when only a tiny percentage of the data is going to be touched, or its promptly going to be paged out again. If he has a machine with 32GB of ram he should retest with a 64GB working set.
That being said, we have an in-memory custom content db where this tweak might really make a difference..
How much exactly do you consider to be the minimum 'enough' RAM where you won't need to page out to disk? 128 exabytes?
I'm not saying 4K pages are good and 1G ones are bad, but there are a fair number of applications that probably benefit from something in between. 1GB is probably on the extreme side of things.
If you have a 50TB database you cannot put it in ram on anything common. The largest machine I've seen for sale takes 16TB of ram http://www-03.ibm.com/systems/power/hardware/795/specs.html.
That doesn't mean you need 50TB of RAM, because 99% of the records could be inactive. Instead you let the hardware page things in, and the portions of the database that are regularly used will stay in RAM, while the rest remains on SSD. The arches with more page selection choices can actually be a big selling point for non x86 servers in certain cases.
In the end, retained data is still growing, and a fair number of applications don't fit into mapreduce or other partitioning schemes. So, I would say paging is going to remain useful for some portion of the servers in existence for at least a few more years.
BTW: I recently worked on an application which would pretty much eat as much RAM (enormous hash table) as it was given and still ask for more. In the end shipping a fairly normal machine (32G-64G) with a 4TB PCIe based SSD provided sufficient performance that we didn't need to spend 100x on a machine that could take 4TB of RAM. So there are economic arguments as well.
Well no kidding, if you're Google then you'd need petabytes of RAM to keep everything in memory.
That doesn't mean the scenario the article was addressing was atypical/unlikely in any way, or that the one you're addressing is more common.
Email is in my profile. Thx!
Ok, from the comments, it's advocating for 1GB pages at the main memory. Not cache, not disk, not network. Main memory.
For me it looks too big - entire servers will have about 32 pages, swapping will take ages at 400MB/s disks. Current PCs use a too small page, but 4MB seems a much more realistic number.
Ultimately it would be handy to be able to tune page size for the loads that you see. I could see page sizes jumping by 4x or 16x (2 bits or 4 bits) each time being reasonable.
The real issue the author is talking about isn't "how much memory does a process allocate" but rather "how many total pages does the OS have to keep track of and what percentage of those fit in the TLB at any one time?"
This hints that there is a tradeoff happening here. At the highest level, the tradeoff is between having efficient use of memory and TLB hits. Big pages give TLB hits, small pages make efficient use of memory.
Since the TLB is in hardware, it is more difficult to have the fine-grain tuning you desire.
But what about virtualization? Can I use 1GB pages in a guest OS, or will the host OS still handle everything with 4k pages, nullifying any advantages?
Short answer: huge pages are a big win for virtualization.
Your application believes that it has all the RAM to itself. This is a lie that the operating system and hardware tell your application to decouple the physical RAM addresses and the ones your application uses (virtual RAM addresses). Learn more about virtual memory here: http://en.wikipedia.org/wiki/Virtual_memory
In order to keep this mirage working, the computer needs to map from virtual address to physical address. Instead of tracking every single address, it tracks spans of addresses. So, the address your application sees as 0 to 4096 will map to physical address 5000 to 9096. Keeping this map using fixed-size spans keeps the size of the mapping down and the performance fast.
This article is about using bigger spans (0 to about 1 billion) instead of the standard 4kb. The advantage of this is that the mapping from virtual to physical is stored in memory as a tree and bigger spans mean you need fewer nodes in the tree. Fewer nodes means you have less traversals/indirection to find the node you are looking for. Less work means faster performance.
The details about the caching and the counts of TLB in the processor has to do with how much dedicated space is in different parts of the CPU for this mapping information.
The details about offsets and changing how the memory was accessed in order get positive / negative performance in the tradeoff of 4kb vs 1gb have to do with wether the mapping information was in the cache or not. it is similar to alignment: http://en.wikipedia.org/wiki/Data_structure_alignment
A lot of the obscure parts of the code are just how the author is calculating addresses to read using pointer arithmatic http://en.wikipedia.org/wiki/Pointer_(computer_programming)#... and bit-shifting http://en.wikipedia.org/wiki/Bitwise_operation
Finally, in order to use these 1gb maps instead of 4kb maps, the programmer has to leverage special way of allocating memory from the operating system called mmap http://en.wikipedia.org/wiki/Mmap
https://github.com/johnj/llds#wall-timings-in-seconds
Same concept applies, reducing translations for page lookups reduces latency.
The site is loading as if the page was 1GB though.