I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.
Under memory pressure, the kernel evicts less popular pages from memory. If a page has been mapped from a file, it is dropped (if dirty, then it is written out first). If it is needed later, the kernel can read it back from the file. If a page is anonymous (read: heap page), then there is no backing file and the kernel copies it to swap before dropping it. This is swapping.
So, what happens if you disable swap and the kernel is low on memory? What can it evict? Anonymous pages cannot be evicted: there is no swap to put a copy in. The only choice the kernel has is to evict pages that are mapped from files. Those include pages mapped from the executable. You don't eliminate stalls by disabling swap, you just move them elsewhere: the kernel will page out code and your app gets paused whenever the execution flow hits such a page.
Your userspace early OOM killer triggers and lets you know you're trying to run more than will fit in memory so you don't do it again. (In my experience, the kernel OOM killer can't be trusted to kill processes soon enough.)
I haven't had to analyze the performance of no-swap processes before. My assumption is that code is hot enough to avoid eviction and that evicted code pages are rather the exception. To strong-man the argument, I can imagine long running complex (bloated) services could have parts that are not touched unless a specific request comes in.
This is not true, or the system wouldn't grind to a halt due to having to reload code pages (even those that are recently in use) from disk over and over.
For example, key presses and output streaming in SSH become unresponsive. The code pages for handling key presses and output streaming are obviously in use — I just used them!
But frankly, as a user I don't really care what the real technical reason is. I just don't think the system should become unresponsive for minutes on end during memory pressure. Just OOM-kill earlier, when it's supposed to.
The proper explanation is relatively straightforward. Normally, any memory request for a process is served by getting a page from the free memory pool. But when the system free memory goes below the low watermark (calculated from vm.min_free_kbytes and vm.watermark_scale_factor), the kernel stops giving the pages from the free memory pool and instead goes into direct reclaim, trying to find unused pages. While the kernel does reclaim, the allocating process remains stalled. The kernel can't return control to the process - there is no page mapped that the process tried to access. This applies to any types of pages - anon, file-backed, code pages. So, under the condition of free memory below the low watermark, any process trying to allocate memory gets stalled. This is measured as memory pressure stalls (PSI) in /proc/pressure/memory.
The problem with the system being stuck in an unresponsive state is that the design choice for the kernel OOM handling is to be optimistic and keep the workload running as long as the kernel is making any progress reclaiming pages. The direct reclaim process needs to fail to make any progress 16 times before the OOM killer is invoked.
The solution is to use cgroup memory limits to avoid getting into the situation of memory exhaustion and use user-space oomd to recover from this situation earlier.
Code pages are paged in on demand: for any sufficiently large binary just some code pages will be mapped right after start. Given enough memory pressure, the active LRU will be shrunk often enough to push code pages out of the active LRU and then they are eventually evicted.
I'm not aware how it works in modern MGLRU, but I would assume similar ovservable behevior holds.
When the kernel needs memory, it goes hunting for a page it can discard. But since that's transparent, the kernel can only discard a page if it knows it can get it back (after all, it's still got valid data on it, and maybe you'll try access it again later).
If there's swap, a page full of stale/unneeded data can be written out to swap. But if there's no swap, your page of "dangling data that you'll never use, but is still valid & referenced" can't be discarded; the kernel doesn't know you won't want it later, and it can't recreate the page if it throws it away.
So like sibling said, at that point it has to find other pages it can evict from memory, ones that _do_ have somewhere persistent they can be written out to. Pages loaded from binaries on disk satisfy that, so those will get dropped instead.
The only code in danger is the one from libraries (e.g. zlib)
You can tell an average Golang user to use mutexes, you shouldn't be telling them to lock memory manually.
If you're in a position where you still want swap to be there, but only used in extreme situations, then you should set it yourself. Otherwise disable swap or switch a language.
Some arbitrary number.
https://docs.kernel.org/admin-guide/sysctl/vm.html#swappines...
> At 0, the kernel will not initiate swap until the amount of free and file-backed pages is less than the high watermark in a zone.
I don't see a reason why the default shouldn't be 0 then on modern systems.
> I don't see a reason why the default shouldn't be 0 then on modern systems.
The zero value makes reclaim very unbalanced and is not appropriate for generic systems. This will cause the kernel to discard all the page cache, slowing down the system with uncached I/O, while keeping all unused anonymous pages in memory.
> If you care about latency, disable swap.
So that only leaves anonymous data pages, i.e., regular data in memory, that could be swapped. Ok but you're not easing the load on code pages, those still get evicted during memory pressure whether you have swap or not.
His reply below was presumably killed by mods for doing this: https://news.ycombinator.com/item?id=49970841
Yes, I was and still running go and zig binaries with mlockall. Apps have large in-application caches and also perform large streaming reads of datasets bigger than available RAM. If it happens that several streaming reads running at same time, they push memory pressure hard enough, so kernel swaps out cold parts of in-app caches. That later leads to elevated latencies => mlockall to avoid that.
This is not exclusive to GC, any program that reads a rarely accessed piece of memory is vulnerable to this.