"Stop doing that"
If you care about latency, disable swap. System wide or for the specific the cgroup.
"Stop doing that"
If you care about latency, disable swap. System wide or for the specific the cgroup.
If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.
I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.
Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.
Under memory pressure, the kernel evicts less popular pages from memory. If a page has been mapped from a file, it is dropped (if dirty, then it is written out first). If it is needed later, the kernel can read it back from the file. If a page is anonymous (read: heap page), then there is no backing file and the kernel copies it to swap before dropping it. This is swapping.
So, what happens if you disable swap and the kernel is low on memory? What can it evict? Anonymous pages cannot be evicted: there is no swap to put a copy in. The only choice the kernel has is to evict pages that are mapped from files. Those include pages mapped from the executable. You don't eliminate stalls by disabling swap, you just move them elsewhere: the kernel will page out code and your app gets paused whenever the execution flow hits such a page.
Your userspace early OOM killer triggers and lets you know you're trying to run more than will fit in memory so you don't do it again. (In my experience, the kernel OOM killer can't be trusted to kill processes soon enough.)
I haven't had to analyze the performance of no-swap processes before. My assumption is that code is hot enough to avoid eviction and that evicted code pages are rather the exception. To strong-man the argument, I can imagine long running complex (bloated) services could have parts that are not touched unless a specific request comes in.
This is not true, or the system wouldn't grind to a halt due to having to reload code pages (even those that are recently in use) from disk over and over.
For example, key presses and output streaming in SSH become unresponsive. The code pages for handling key presses and output streaming are obviously in use — I just used them!
But frankly, as a user I don't really care what the real technical reason is. I just don't think the system should become unresponsive for minutes on end during memory pressure. Just OOM-kill earlier, when it's supposed to.
The proper explanation is relatively straightforward. Normally, any memory request for a process is served by getting a page from the free memory pool. But when the system free memory goes below the low watermark (calculated from vm.min_free_kbytes and vm.watermark_scale_factor), the kernel stops giving the pages from the free memory pool and instead goes into direct reclaim, trying to find unused pages. While the kernel does reclaim, the allocating process remains stalled. The kernel can't return control to the process - there is no page mapped that the process tried to access. This applies to any types of pages - anon, file-backed, code pages. So, under the condition of free memory below the low watermark, any process trying to allocate memory gets stalled. This is measured as memory pressure stalls (PSI) in /proc/pressure/memory.
The problem with the system being stuck in an unresponsive state is that the design choice for the kernel OOM handling is to be optimistic and keep the workload running as long as the kernel is making any progress reclaiming pages. The direct reclaim process needs to fail to make any progress 16 times before the OOM killer is invoked.
The solution is to use cgroup memory limits to avoid getting into the situation of memory exhaustion and use user-space oomd to recover from this situation earlier.
Code pages are paged in on demand: for any sufficiently large binary just some code pages will be mapped right after start. Given enough memory pressure, the active LRU will be shrunk often enough to push code pages out of the active LRU and then they are eventually evicted.
I'm not aware how it works in modern MGLRU, but I would assume similar ovservable behevior holds.
When the kernel needs memory, it goes hunting for a page it can discard. But since that's transparent, the kernel can only discard a page if it knows it can get it back (after all, it's still got valid data on it, and maybe you'll try access it again later).
If there's swap, a page full of stale/unneeded data can be written out to swap. But if there's no swap, your page of "dangling data that you'll never use, but is still valid & referenced" can't be discarded; the kernel doesn't know you won't want it later, and it can't recreate the page if it throws it away.
So like sibling said, at that point it has to find other pages it can evict from memory, ones that _do_ have somewhere persistent they can be written out to. Pages loaded from binaries on disk satisfy that, so those will get dropped instead.
The only code in danger is the one from libraries (e.g. zlib)
You can tell an average Golang user to use mutexes, you shouldn't be telling them to lock memory manually.
If you're in a position where you still want swap to be there, but only used in extreme situations, then you should set it yourself. Otherwise disable swap or switch a language.
Some arbitrary number.
https://docs.kernel.org/admin-guide/sysctl/vm.html#swappines...
> At 0, the kernel will not initiate swap until the amount of free and file-backed pages is less than the high watermark in a zone.
I don't see a reason why the default shouldn't be 0 then on modern systems.
> I don't see a reason why the default shouldn't be 0 then on modern systems.
The zero value makes reclaim very unbalanced and is not appropriate for generic systems. This will cause the kernel to discard all the page cache, slowing down the system with uncached I/O, while keeping all unused anonymous pages in memory.
> If you care about latency, disable swap.
So that only leaves anonymous data pages, i.e., regular data in memory, that could be swapped. Ok but you're not easing the load on code pages, those still get evicted during memory pressure whether you have swap or not.
His reply below was presumably killed by mods for doing this: https://news.ycombinator.com/item?id=49970841
Yes, I was and still running go and zig binaries with mlockall. Apps have large in-application caches and also perform large streaming reads of datasets bigger than available RAM. If it happens that several streaming reads running at same time, they push memory pressure hard enough, so kernel swaps out cold parts of in-app caches. That later leads to elevated latencies => mlockall to avoid that.
This is not exclusive to GC, any program that reads a rarely accessed piece of memory is vulnerable to this.
To be more explicit about my advice: disable swap and ensure you have sufficient memory for your workload.
Your 14 MB Go binary is not the source of your memory pressure. Your memory pressure is coming from the allocated anonymous pages supporting your programs data structures on the heap.
Have sufficient memory for your workload, monitor it, and let the Go GC and kernel manage memory normally. This will work exactly as you'd hope.
You are correct, that 90MB Go binary is not major memory consumer. But it doesn't matter what is the source of memory pressure. What matters is what kernel do under pressure. If you disable swap, then under memory pressure kernel can only pageout memory mapped files, no matter how small they are. And your code will be paged out.
Introducing mlock is operationally annoying and trying to solve the wrong problem. It requires elevated container permissions, it's not something your SRE team is going to expect, few programs do it, and it has all kinds of technical implications and interactions with kernel OOM system, forking, the Go GC, etc.
And nothing about mlock solves the problem of having insufficient RAM for your workload.
There are very specific scenarios where mlock is exactly the right solution but the typical Go HTTP service is not one of them. The KISS principle applies.
The good thing is that mlock/mlockall does not require elevated permissions. But it is bounded by the memlock rlimit, which you need to configure for the container. If you don't want paging/swapping, set it equal to the total memory you give to the container and call it a day.
I'm curious about the technical implications. Which ones do you have in mind?
To be honest, I don't quite see how mlock makes it complex. It's quite the opposite: it has very clear semantics and makes reasoning about system behavior simple. Disabling swap is the opposite: you're making a bet and hoping it works.
Have you actually done this yourself for Go programs that you've deployed into production? If so, what were the circumstances? Was it your solution to latency spikes?
The alternative (which I’ve also done), is to put the binary in tmpfs, after making sure it’s not bloated.
Or did someone else do this work and you assumed it was effective? Or did you personally implement this change and then validate that the before/after metrics improved as a result?
The impact is pretty obvious.
Just seems kinda weird to give advice on forums when you're not speaking from actual relevant experience. I'm getting meat proxy answers and confident replies from people who've never done what they're suggesting others do.
I ultimately change the runtime environment to no longer allow any local durable storage for job (with a handful of exceptions), including swap and binary staging. The fact that almost everything was already using mlock made the transition simpler.
Java is an example of the most widely deployed garbage-collected language in production around the world, deployment number and scale-wise. Most of the world's missing critical systems today run on Java. Java can run as a process or in a container. Containerised Java deployments do not require unique operational SRE arrangements, nor do they require any unique/dedicated permissions to run in a container.
Upon start-up, the JVM checks whether the system has sufficient memory to satisfy the max heap size constraint, and it will duly fail to start up if not. Afterwards, the JVM will allocate the required amount of memory, and it will mlock it. Multi-terabyte[0] RAM JVM deployments[1] do exist and exhibit exactly the same behaviour. Java has had such behaviour since its inception.
The prevailing number of Java deployments run on systems with swap enabled, and, from the operational standpoint, it is pure insanity to run mission critical (banking, payments, government etc) apps without a swap. Swapless JVM deployments are virtually unknown or are operational negligence.
The principal exceptions for running a JVM on a system without a swap are hard real-time scenarios and embedded systems. That is it.
What makes Go different or unique that it can't use mlock and that it would require special permissions or unique/distinct SRE/operational arrangements? It is eluding me.
[0] The Azul JVM supports up to 20 Tb of RAM which it will mlock.
[1] See the published Azul/Forrester evidence for reference. It is an interesting read regardless of the mlock debate.
Java is generally simpler to run without swap since you know in advance how large the heap can be, so it's just a small amount of variability in the jvm internals to consider.
You don't want swap of heap/stack, or paged out code pages. The latency hit is pretty bad even for non-"real"-time systems. And really, you should know how much memory you are going to use (to within some fair bounds).
Adding only a small amount of swap for very cold pages does not save very much ram. And using it for anything else is very bad for interactive systems. In the same vein paging out code pages you won't ever run just does not save very much.
It requires some discipline in monitoring, load testing, etc, but you should be doing that anyway.