Disable transparent hugepages
blog.nelhage.com
blog.nelhage.com
[1] https://www.cs.rice.edu/~druschel/publications/superpages.pd...
Given alc@ was an author on the paper (and the paper's FreeBSD 4.x implementation supported multiple superpage sizes), I'm not really sure why FreeBSD's pmap doesn't have support for 1GB page promotions.
The only mailing list thread I see regarding this is here, and it doesn't seem particularly underhanded to me: https://lists.freebsd.org/pipermail/freebsd-hackers/2014-Nov...
https://lists.freebsd.org/pipermail/freebsd-hackers/2014-Nov...
But maybe I'm missing something?
See also: https://lists.freebsd.org/pipermail/freebsd-hackers/2013-Sep...
There's been some work to improve performance (e.g. https://github.com/torvalds/linux/commit/7cf91a98e607c2f935d... in 4.6) but I haven't tried if this fixes my workload.
https://access.redhat.com/documentation/en-us/red_hat_enterp...
https://alexandrnikitin.github.io/blog/transparent-hugepages...
Especially the conclusion is noteworthy:
> Do not blindly follow any recommendation on the Internet, please! Measure, measure and measure again!
It baffles me that THP became enabled by default (is it? I think it’s only a default on RHEL distros?). It really screws up many expectations that applications might assume about memory behavior (like the page size). In the majority of cases, THP is a bad, bad idea and anyone with perf or devops experience will agree with this I think.
Do you want to impose GC like pause characteristics to all processes on your box? And possibly double, triple, or 10x your memory usage? Enable THP then.
I can confirm that neither my Ubuntu nor Debian servers have is in "always", they're either "madvise" or "never".
# CONFIG_TRANSPARENT_HUGEPAGE_ALWAYS is not set
CONFIG_TRANSPARENT_HUGEPAGE_MADVISE=y
and /boot/config-4.13.0-1-amd64 has: CONFIG_TRANSPARENT_HUGEPAGE_ALWAYS=y
# CONFIG_TRANSPARENT_HUGEPAGE_MADVISE is not set
So this is a recent change.Edit: The linux kernel source says the default is always (in mm/Kconfig), and that's been true since 2011.
The debian package changelog says the change occurred in 4.13.4-1:
* thp: Enable TRANSPARENT_HUGEPAGE_ALWAYS instead of
TRANSPARENT_HUGEPAGE_MADVISE
The reason is not given in the changelog itself, but it's given in the git log of the debian packaging:As advised by Andrea Arcangeli - since commit 444eb2a449ef "mm: thp: set THP defrag by default to madvise and add a stall-free defrag option" this will generally be best for performance.
https://anonscm.debian.org/cgit/kernel/linux.git/commit/debi...
Edit 2: The mentioned commit (444eb2a449ef) dates back to 4.6, so presumably, at least some performance issues with transparent huge pages may be gone since that version of the kernel.
$ cat /sys/kernel/mm/transparent_hugepage/enabled
[always] madvise never
I can also confirm that my Ubuntu 16.04.3 LTS (GNU/Linux 4.4.0-101-generic x86_64) droplet has the same setting.
Makes me think that your setting is a default and his was set by Digital Ocean.
Perhaps there are issues with 'madvise' in kernels prior to 4.11, so they chose 'always' rather than 'never'.
It is in the process of being fixed (it should be set to madvise).
Since "enabled=always" is the kernel default value, anything that uses a stock kernel (example: Arch family) or has to build its own (Gentoo) will probably have it enabled by default.
I just checked, and my Gentoo and Manjaro systems have it set to "enabled=always".
Anyway, recently they added new "defer" mode for defragmentation so THP doesn't try to defragment (the main cause of the slow down) upon allocation and instead it is triggering it in background (via kswapd and kcompactd). This is now set to be the default. I think it is available in RedHat/CentOS 7.3+
The best out of both worlds though (although then it requires more manual work) is pre-allocating HugaPages in advance and then let application use them (if the application supports it) or through libhugetlbfs (if it doesn't).
Edit: changed hugetlbfs to libhugetlbfs, so it's easier to find how to do it with man libhugetlbfs
Of course you could profile and measure performance to determine if the warning is applicable but is that something I should be doing for every part of the stack? I should but should I prioritize that over x, y or z?
So apparently, transparent hugepages have some issues in their current implementation that can cause big performance losses in some cases. Seems to me like that's a bug, and I see no reason why that bug couldn't be fixed in the future.
By following random recommendations, you get into situations where the underlying problem has been fixed for ages, but people still cargo-cult some workaround that actually makes things worse with the new implementation.
And it doesn't print a message like "yeah I stalled your box for the last 60 seconds in order to shuffle deckchairs around, sorry" in syslog.
So you pull your hair out trying to figure out why your nice stable service all of a sudden sets off Nagios at 2am for no obvious reason, every week or two.
Similar to the issues you had with Redis, the kernel change to THP on by default totally destroyed CoW sharing for forked Ruby processes, despite Koichi Sasada's change to make the GC more CoW friendly. Without disabling THP, a single run of GC marking can cause the entire heap to be copied for the child.
It's also important to measure in your actual use-case, and not just with benchmarks that seem "close enough"; I know it sounds odd, but I've seen others adjust settings and then prove that it worked with a benchmark that they claim is "representative", when in reality they didn't actually improve anything because that "representative benchmark" differed from the real use case in precisely the way that would not respond to the adjustment.
Blindly following "best practices" is bad enough, but "proving" that the changes work with crucially-different benchmarks is worse; and when it's some expensive consultant doing such things, I think it may even approach fraud.
That is exactly the reason I wrote the post! Those advice are based on specific use case, bug or outdated kernel. The jemalloc (Digital Ocean post) case is a good example, it just doesn't (didn't) know about THP https://github.com/jemalloc/jemalloc/issues/243
I can only repeat it: "Measure, measure and measure again!"
(I don't want to preallocate hugepages because KVM is only a small part of my workload.)
At the lower syscall level you move the BRK address, which is the highest memory address you're allowed to use. By default this is just after the statically initialized memory.
malloc() is just a library that manages this memory for you.
Linux has no idea if you will use the memory you just allocated, usually this happens dynamically; when you access a memory region for the first time, it is allocated in memory for real.
For example, if I map a bunch of pages of memory, use some, and then set MADV_DONTNEED on just a few of those pages afterwords (so they can be given back to the system temporarily), this will only work if I know the entire page is unused. If a page magically gets 512x larger under my feet, it's possible it will "coalesce" with another page -- into a huge page -- that can't be given back, because some of the (now much larger) page is still needed. This is the case of what happens with `jemalloc` + `redis`, where what looks like a memory leak is actually a failure to give back pages to the operating system, because small pages coalesce into huge ones automatically, defeating MADV_DONTNEED.
Whether or not malloc(2) uses THP "under the hood" in this case is more an implementation decision (there are many malloc variants), but it just punts the problem down the road. Ultimately it manifests as a violation of the "unspoken protocol" between the kernel and application when managing memory mappings.
All that said, there's some posts elsewhere in this thread pointing out that FreeBSD's Huge Page implementation avoids some of these deficiencies (possibly as a trade for some of its own), so all that said -- there's still room for better implementations, clearly: https://www.cs.rice.edu/~druschel/publications/superpages.pd...
The issue happens on specific workloads (databases, hadoop etc) and (this is often not mentioned) after the system is running uninterrupted for a quite of while. The slow down comes due that the workloads mentioned cause memory to be fragmented and when kernel tries to defragment the memory (unsuccessfully) on each allocation.
Since the workload you mentioned looks like it is for a workstation that won't be running a database 24/7 over months/years, you are very unlikely to run into it.