Open sourcing oomd, a new approach to handling OOMs
code.fb.com
code.fb.com
Printing this seems like a good addition to "Linux Performance Analysis in 60,000 Milliseconds". [2] More useful that the load average printed by uptime, which hasn't aged well. Linux mixes cpu-bound and io-bound stuff into one bucket [3] which is a terrible idea, and the number of tasks is confusing at best when you have a mix of thread-per-CPU and thread-per-connection stuff running.
Looks like it's also exported on a per-cgroups basis, which is great.
[1] http://git.cmpxchg.org/cgit.cgi/linux-psi.git/tree/Documenta...
[2] https://medium.com/netflix-techblog/linux-performance-analys...
[3] http://www.brendangregg.com/blog/2017-08-08/linux-load-avera...
I'm sure I'll use this once there's OSes running a stable kernel with these patches, but I don't run non-vendor kernels in production [unless I'm forced to], and depending on kernel features makes it hard to support legacy systems.
edit: you seem to be implying that FB is stupid and likes writing kernel code for no good reason.
I prefer to set vm.overcommit_ratio=0 and then increase vm.vfs_cache_pressure to somewhere between 400 and 10000 depending on the server role, then set vm.min_free_kbytes, vm.admin_reserve_kbytes and vm.user_reserve_kbytes higher based on the amount of ram in the system using a simple formula.
For systems that are ephemeral, I also set vm.panic_on_oom to 2 so that they self heal. Avoid doing this on databases.
Of course, THP should also be set to madvise and memory defrag disabled unless you really need it. That will also free up some leaky memory and artificial pausing.
Beyond these things, I agree that CGroups is another way to set memory constraints around applications to further protect the server. There are many that won't have CGroups for a while however, as many systems are still on older non systemd distributions.
For my curiosity, is there a technical reason for the distro to use an old kernel? I would think one would wish to have all the bug fixes available, and I don't think the kernel has a bad backward compatibility story, so from the outside it seems like a sure thing -- leading me to believe there must be more to the story.
I understand vm.overcommit_ratio=0 but can you elaborate on what the net effect of setting all the rest of of these tuneables to these values is? What does this achieve and how do these work together?
>"For systems that are ephemeral, I also set vm.panic_on_oom to 2 so that they self heal. Avoid doing this on databases."
I'm guessing this triggers a reboot for you when the OOM evicts a process? Do you do this in conjunction with the above setting? Or is this an alternative?
[1] - https://www.kernel.org/doc/Documentation/sysctl/vm.txt
All of the settings should be calculated per server role and memory capacity after extensive testing of your application settings and performance.
> I'm guessing this triggers a reboot for you when the OOM evicts a process?
It triggers a panic if you hit OOM, rather than kicking in OOM killer which often leaves a system broken, sticky and confused. This should only be used for ephemeral systems or systems that do batch work and can handle reboots. This requires people to find the root cause of their over-allocation and fix them.
So you're basically turning every node/instance into a special snowflake of sysctl settings all in the hopes of preventing an OOM? And you do this every time dev push a new rev of code? This doesn't seem very agile or sustainable at any of kind of scale or release velocity.
Keep an eye out for patches making their way to distribution kernels around THP / memory. There has been some major work done upstream and THP is no longer the alloc-stall causing monster that it was before. It used to cripple some fleets around here (and so we'd have it completely disabled), but once Oracle backported the fixes in to UEK4 it has been smooth as silk with it turned back on.
If you're interested in experimenting and seeing if the patches fix any THP problems you see, UEK4 is completely free, and can be installed on RHEL7/CentOS7 etc. (http://public-yum.oracle.com/repo/OracleLinux/OL7/UEKR4/x86_...).
Could maybe use it to provide details in a bugzilla report?
Similarly for disk space being low, or under IO pressure, some apps could hold off checking certain files / throttle things for a while.
If your app has a cached process and it retains memory that it currently does not need, then your app—even while the user is not using it— affects the system's overall performance. As the system runs low on memory, it kills processes in the LRU cache beginning with the process least recently used. The system also accounts for processes that hold onto the most memory and can terminate them to free up RAM.
0% of applications correctly handle being unable to write to a page they were given as writable. It's certainly true that few applications correctly handle a failed malloc, and in many cases there's not really a correct way to handle it, but at least it's available to handle.
- introspection at the application level: exit if memory usage goes way above usual (say by 2x)
- management by systemd to gracefully restart the application when it exits (dependancies etc), and also monitor memory use in case the application doesn't monitor it well (say exit when 2.5x the usual)
I could be made better, but at the cost of more complexity.
To give more context: I run many copies of a daemon with known memory leaks, on machines each with a load close to 0.7 and about 200M of free ram left. I consider that an heavy load.
This approach restarts one of the copies of the daemon when its own leaks become dangerous to the rest of the system.
It is not perfect, but good enough when the numbers are measured at peak. The 200Mb free RAM are all the extra leeway I care to leave.
After playing around with vm.overcommit_ratio, different swap sizes, earlyoom[1], and a few other variables, I still haven't found a happy medium between high memory utilization and low risk of swapping to death. vm.overcommit_ratio=0 is safe, but on systems where occasional swapping is tolerable and memory is limited (e.g. my laptop), I'd rather allow some overcommit.
The risk is that if many cold or unallocated pages get touched while the system is under high memory pressure, the system can become totally unresponsive. At the moment I use "Magic SysRq"+f to manually start the oom_killer, when possible. Obviously it's not a great solution. Is there some kernel tunable to keep the system responsive that I'm unaware of? What do you guys do for desktop/laptop systems?
The other aspect of your question is about how "many of us" are working at places with huge datacenters. If you count up all the folks at Amazon, Google, Microsoft, Facebook, and Apple how does that compare to the long tail of people operating smaller datacenters at smaller companies?