I've had bad luck with transparent hugepages on my Linux machines
utcc.utoronto.ca
utcc.utoronto.ca
I did a lot of HPC modelling ranging from hundreds of GiB to TiB-sized RAM servers and THP was an instantaneous win over not using it. Later on I experimented with LD_PRELOAD_PATH and libhugetlbfs and while it did stabilise things even more and reduced time spent in the page table, it was not even several factors 'better' than THP.
A large part of this really boils down to your performance needs. Deterministic modelling -> fairly stable and reproducible malloc patterns -> THP will probably work OK.
If your memory usage spikes a lot then THP is probably more hassle than it's worth. The kernel will spend an eternity thrashing and kicking the page table and trying in vain to clean up after itself. I can see why people feel THP sucks under those conditions!
If anything, the real crime here is how awful Linux is at HPC without a lot of tuning and careful tweaking. Throw NUMA into the mix, file caches and the OOMkiller and it feels like it never moved out of the 90s. Combine it with poorly-configured hypervisors and your performance will yo-yo and you'll spend an eternity trying to figure out why (ask me how I know...)
the real crime here is how awful Linux is at HPC without a lot of tuning and careful tweaking
To what are you making the comparison? Are *BSDs better? HP-Unix? AIX? Windows? What's the competition in this space that makes Linux look bad in HPC?Edit: Added "in HPC" at the end for clarification.
Even the workloads that use GNU/Linux, most likely are heavily customized versions provided by IBM, HP and co.
https://www.ibm.com/high-performance-computing
https://www.hpe.com/us/en/compute/hpc.html
I also remember that for a while IBM's xl compilers were used quite often, not sure if that is still the case.
I don't know exactly what the current Cray environment is like -- it used to be rather odd -- but most HPC systems run normal EL-ish distributions. Summit is RHEL8, and mostly uses GCC, not XL according to https://gcc.gnu.org/wiki/cauldron2022talks?action=AttachFile...
This makes sense to me
If OOMkiller kicks in it suggests a resource manager issue (not accounting memory). Then, you can't blame the kernel for NUMA, which you presumably want for performance; I don't understand the difficulty people have with pinning processes after many years of it being necessary. There's something to be said for userspace filesystems too, like the venerable PVFS. The answer to hypervisors should be don't do that.
In fact it's the _absence_ of choice that hurts more than it is a surplus of choice.
As for OOMkiller: I think we're both wise enough to know that experimentation is a large part of what any team that consumes gobs of RAM do. So talk of a "resource manager issue" is all well and good in prod when you should have a reasonable handle on what's used by what and for how long. Less so when you're scaling a model --- say a monte carlo model -- to a larger number of simulations in development and you're testing things. When you're paying an awful lot of money for hardware you start to count (or you should, if you respect your company's money) the costs of things like this.
Regardless, the OOMKiller will indeed reap stuff; and more often than not, it'll pick something shouldn't (it's probabilistic and wrong as much as it is right) and that can cause headaches.
As for NUMA: I'm not blaming the kernel for anything :-)
And hypervisors are sometimes a given, and not a choice. We don't all get to pick our hardware, nor what hardware is made available to us.
Actually, I was forgetting the stupidity of the memory cgroup invoking the OOMkiller, rather than giving ENOMEM, if you purely use the cgroup for memory accounting -- a sensible-looking change from OpenVZ was rejected. That means there's probably no indication about what's happened unless the job is correlated with the syslog message. As I don't get to do that stuff any more, I haven't looked for a way to hook it now. There's obviously a window in which it can fail, but the resource manager can track the job's PSS, as well as at least ulimiting data and stack. You certainly don't want the job to start paging, if you have swap -- that definitely wastes the resource. If you really don't want to limit jobs' memory use, the resource manager can at least adjust the oom_score.
Note glibc lets you turn on and off THP per process which is pretty useful for benchmarking if it helps or hinders performance.
$ hyperfine ' nbdkit -U - data "1 * 10737418240" --run exit '
Benchmark 1: nbdkit -U - data "1 * 10737418240" --run exit
Time (mean ± σ): 3.658 s ± 0.049 s [User: 0.406 s, System: 3.242 s]
Range (min … max): 3.576 s … 3.713 s 10 runs
$ hyperfine ' GLIBC_TUNABLES=glibc.malloc.hugetlb=1 nbdkit -U - data "1 * 10 737418240" --run exit '
Benchmark 1: GLIBC_TUNABLES=glibc.malloc.hugetlb=1 nbdkit -U - data "1 * 10 737418240" --run exit
Time (mean ± σ): 1.655 s ± 0.007 s [User: 0.299 s, System: 1.350 s]
Range (min … max): 1.643 s … 1.666 s 10 runsBut there was a surprise... more than 10 times degradation of overall Linux server performance due to increased physical memory fragmentation after a few days in production: https://github.com/ClickHouse/ClickHouse/commit/60054d177c8b...
It was seven years ago, and I hope that the Linux kernel has been improved. I will need to try "revert of revert" of this commit. These changes cannot be tested by microbenchmarks, and only production usage can show their actual impact.
Also, we successfully use huge pages for text section of the executable, and it is beneficial for the stability of performance benchmarks due to lowering the number of iTLB misses.
[1] ClickHouse - high-performance OLAP DBMS: https://github.com/ClickHouse/ClickHouse/
I have only the most basic knowledge of hugepages - picked up on it being a necessity for stable performance in VMs, but don't understand why it would matter for Postgres as well.
And do you end up having to configure Postgres to use hugepages, or does it pick them up automatically?
Postgresql needs to be reconfigured to use them properly. The shared_buffers setting needs to be adjusted to allocate in [huge page size] units.
@menaerus actually has a great comment on this that digs deeper into it, looking forward to seeing what comes out of that thread as well.
hugepagesz=1G hugepages=940
and then add these config lines to postgresql.conf: huge_pages = on
huge_page_size = 1GBNot coincidentally, Microsoft's primary use case for 1GB pages on Windows is MS SQL server.
It is automatically enabled, but only if you have the "Lock pages in memory" privilege assigned to the database engine service account...
... which the default SYSTEM account has enabled by default...
... but not if you set up a typical cluster with a domain user account.
So in other words, the scenario where you would want this enabled the most -- large expensive enterprise clusters -- is where it is accidentally disabled with no warning.
I use this as a free 10-20% performance boost. Add in a few other similar tuning settings and I can get a 50% pref improvement on just about any "enterprise" cluster without having to get clever.
Turbo buttons are fun to press.
Have you ever written up some of the things you do to improve performance, or is that something you're unable to share?
https://www.sqlshack.com/category/sql-server-performance-tun...
- Enable the "Lock pages in memory" privilege.
- Enable the "Perform volume maintenance" privilege.
- Enable the "Ad hoc query optimisation" setting.
- If on-prem, set power management to "High Performance", ideally in the hardware BIOS, but failing that, at the hypervisor level.
- Upgrade the OS and DB engine version if this is compatible with the apps.
- Run sp_Blitz and implement any critical recommendations.
The below recommendations apply to cloud-hosted servers only:
- Upgrade to the latest gen DB-optimised VM sizes. In Azure I use Ebds_v5 and Eads_v5.
- Use the "temporary storage" SSD volume for tempdb.
- Consolidate storage. Striped a single large logical volume across multiple 1 TB disks. NEVER use the 1990s era many-small-volumes scheme of having separate data, log, tempdb, and templog drives!
- Move the DB VM and the App VMs into a single "Proximity Placement Group" or the equivalent construct.
In my experience the above list generally yields a total performance improvement of anywhere from 40% to 300% with minimal risk.
If you sprinkle some light tuning on top, such as judiciously creating a handful of indexes with a high predicted performance benefit can take you even further.
Past that, you have to get elbow deep into the schema, which is typically only possible with "in house developed" databases.
That's quite interesting. How big the page-table must be on your server to directly or indirectly cause an OOM?
4-level page-table data-structure can address 512^4 page-tables. A single page-table can contain 512 entries. And each page-table entry is 64 bytes large. Unless I missed something, this means that upper bound memory consumption of page-table data-structure is 512^4 * 64B and which is 4TB of RAM so I guess it is theoretically possible to OOM on a 1TB machine.
What I don't understand is how huge-pages would help you mitigate this problem because 2MB huge-page entry will essentially be made up from 512 entries from a single table. Linux call them compound pages. And given that all those entries will be vacant, and in-use for that particular huge-page, upper bound memory consumption will remain the same as with 4K pages.
It has been recognized that compound pages might be wasteful because all those 511 entries are basically pointing to a head entry [1], so unless you're running recentish version of kernel on your system, you wouldn't see this advantage.
So, I have about 95% of the RAM assigned to postgres shared buffers. Without hugepages enabled, this was fine when the database started up, but as (presumably) the buffer cache became fragmented over time, the size of the page tables grew to a point where the machine ran out of RAM. I do have vm.overcommit_memory = 2 (never overcommit), but even so the OOM killer was still invoked because, well, the machine had no RAM left. I can't remember exactly how big the pagetables had gotten in this situation, but I think it was a few tens of GBs.
I'm actually using 1GB hugepages, but in all honesty I'm not sure how Linux deals with them or how postgres claims them. From /proc/meminfo it looks like the page table is quite small, with the "Hugetlb" and "DirectMap1G" numbers both around 1TB.
I, too, find the file cache to be significant performance impediment on high-RAM systems. Worse, its behaviour is non-deterministic (or at least hard to reason about.)
It leads to page table fragmentation and that is the ultimate performance destroyer. Huge page tables are a godsend in this case, but you do need a program that is properly optimised to handle them properly. At least most RDBMS do this reasonably well.
Mitigating the OOM by switching to hugepages may suggest that you're running a kernel with the optimized page-table hugepage handling because of which there are less page-table entries and consequently the whole page-table ends up being smaller. Currently, I have no other explanation.
The problem with the page table and postgres is that they're separately maintained for each process, even if all processes share the same shared memory area.
With 1GB or 2MB pages that's not too bad. But with 4k pages you can easily end up with dozens to hundreds of MB for each connection. Wasted memory + wasted cycles (tlb misses).
It was requested by the Oracle DB team
https://lwn.net/Articles/919143/
I recognize your username so you'd know better than most whether Postgres would be able to use this or not though
I'd expect it to benefit postgres substantially.
I am eagerly awaiting its release as a DB enthusiast and lifelong Postgres user
That would make sense, in retrospect: the OOMs almost always occurred at the point of highest client load.
However, you're right that multiple processes maintaining separate page-tables where each is containing mappings to the same physical memory regions can be wasteful.
Still, due to how huge pages are represented in the page-table data structure, I don't see what difference would it make to switch to 2MB or 1GB page, unless you're running a recentish kernel (~3 years) that has been specifically optimized for exact this case, e.g. to reduce the memory footprint of huge-page representation in page-tables. Specifically, a single 2MB huge page is now represented by 8 page entries (512 bytes) instead of 512 page entries (2MB). And there was another patch proposing to cut 8 page entries to only 2 page entries.
In any case, that's a dramatic cut of memory footprint when compared to how it used be and which is why I believe that this is the actual reason why OP sees the difference.
My "tune for vmware" script has grown to three lines over the years:
echo never > /sys/kernel/mm/transparent_hugepage/defrag
echo 0 > /sys/kernel/mm/transparent_hugepage/khugepaged/defrag
echo 1 > /proc/sys/vm/compaction_proactiveness # default 204 kB pages reduce CPU performance due to frequent TLB misses and make virtual memory perform worse due to number of faults.
Use 1G pages, and boom, TLB misses and faults are just gone. But definitely not a good default choice!
Any choice is a tradeoff. The larger the page size, the more physical memory is going to be wasted. On the other hand, I guess with the current NVMe hardware the effect on swap might not be too bad.
64 kB pages would mean the number of TLB misses and faults are just one sixteenth of the current status quo. Diminishing returns and all.
4k/2MiB/1GiB
16KiB/32MiB
64KiB/512MiB/4TiB
That said, enough useful "workstation" things assume 4K and break on 64K that I'm hopeful aarch64's use of 16K pages will get people to realize everything doesn't use the same page size. I see transparent hugepages as trying to have your cake and eat it too.
I believe the consensus is that Workstation makes lots of memory allocation and fragment the memory, causing the defragment-er to kick in.
Vmware says won't fix lol.
And everything makes sense now. I was having terrible terrible problems with Java and transparent huge pages. It only affected Java. I think this is because Java was the only thing we were running that would pre-allocate a large chunk of address space for the heap like this. I ended up disabling THP.
However, on a more recent install, the problem magically went away, and I got the impression that the problem was fixed. The article is recently-written, but they're not using an old Linux install, are they?
In future cases it would be useful to look at the stats in /sys to see what it is doing. Under non-OOM conditions it’s not easy for it to use a lot of CPU because it defaults to scanning only small runs of pages and waiting ten seconds between scans.
On a stock kernel, it's 511. TCMalloc's docs recommend using max_ptes_none set to 0 for this reason: https://github.com/google/tcmalloc/blob/master/docs/tuning.m...
(Disclosure: I work on TCMalloc and authored the above doc.)
But to get back on topic, I’ve also had some bad luck with these but that was mostly due to me making the mistake of using them for a database (I didn’t RTFM).
Is that the basic problem?