- swap on spinning rust outright kills the system if memory is low. it's a full blown crash.
- swap on ssd act's much the same
- zram swap stalls and slows down the system for several minutes
I just raise vm.min_free_kbytes to make the oom killer faster and disable swap. At least the machines don't crash hard this way.
Me and my team removed swap on tens of thousands of machines cause we were sick of dealing with this. We wanted the machines to fail hard and fast and not go into a state where it’s doing only a minuscule amount of actual work while trying to recover.
On most production systems, memory allocation is roughly 1% system and 99% application (not counting temporary, evictable allocations like disk cache). Modern applications are not designed to have their pages swapped out to disk and it is not particularly helpful to swap the teeny tiny bits of OS.
There are parts of this article which are true, but the overall picture it paints is not accurate and the conclusion that swap is helpful is not correct for high performance systems - both workstations and servers.
(The other argument is the usual mix of "but it can swap out unused areas" in which case you should fix the program to not allocate a ton of memory it doesn't need; and "but the OOM killer will come" and yes that's the entire point please move my process elsewhere).
Isn't this basically asking every application to independently reimplement their own swap-style functionality? Surely it makes more sense to have swap as a system level feature that every application can take advantage of?
That way the system has more information available to make swapping decisions since it can take into account the memory usage of all applications and drivers/etc together, and applications can take advantage of the most advanced swapping algorithms in the latest versions of the system without having to be individually updated.
No, it's asking for applications not to allocate a bunch of space they don't need in the first place.
In practice I expect the typical scenario you're thinking of is applications are allocating a bunch of memory which they use only briefly or rarely and could easily fetch or recompute that data again later if needed. That's exactly the scenario which swap optimizes for. So why should every application individually implement logic to optimize for it? Fixing these kind of issues is not as simple as "just don't allocate the memory".
That is what I'm suggesting, and I suggest you look at how much space is wasted by startup-initialized data in libraries for features you'll never use, JIT representations that never get compiled for more than one shape, class metadata that never gets touched and vtables which don't get used, how easy it is in any GC language to accidentally keep a large buffer alive until the end of a lexical scope rather than its last real use, etc. etc.
Today's programs are bloated as hell dude.
Especially if it’s a library or the like?
That infrequent call might be key for program stability in one context or pointless (garbage collector whatever- key in a major server program, pointless in a toy app).
The moral of the story is, if under load code or data gets moved from where access is fast (RAM) to where it is slow ( even SSD and most NVMe counts here) then gets accessed when under load? It makes the problem much worse.
Hence why the advice for servers is turn swap off. It’s a footgun waiting to happen.
Once you exhaust the RAM and you run with no swap configured, what happens with your application?
That rarely used code path will now take 50 seconds because it got called when it was swapped out (and the system is under memory pressure), instead of either OOM’ing awhile ago or completing in a couple microseconds like it usually does.
And ‘infrequently called’ here could be every couple minutes.
I had it happen recently that the backup/sync software for a NAS had a constant factor memory consumption based on the number of files it was syncing. Transferred in a bunch of data, and blam. Commercial NAS, and they used swap (shitty, never buy QNAP). Wedged so hard it took a hard power cycle to even get console, AND caused data corruption in the ZFS pools.
It’s what happens next that decides the stability of the system.
If it can grow into swap and keep going, things start crawling, load builds up, buffers expand, and the system eventually grinds to a halt. On Linux, often in a really wedged and irritating/impossible to fix way.
Or, if no swap, OOM killer shoots something in the head (hopefully the offender), and we’re back to normal (minus the thing it killed, which is usually the thing you didn’t want it to - but at least the system works and you know something is wrong).
In either case, caches have dropped to near zero awhile ago so performance is already getting bad. It’s if we enter a death spiral, or get death early enough the whole system doesn’t spiral.
Right. You can get pretty deep into a death spiral even without swap. I feel like the better-performing solution is to make OOM trigger earlier, and to then go ahead and have swap be on.
Somewhat tangentially, I'd really like a setting for minimum disk cache, and that would do so much to help prevent thrashing.
No, swap makes it worse. In virtually all cases I’d rather die than stall or flap.
I’m sure someone here will chime in with their example though.
I rarely need or set up much swap on a server, but I've managed to have the same kind of problem with the OOM killer not kicking in very quickly on a server. And on my desktops swap is able to soak up many gigabytes and improve my performance a lot with thrashing basically never happening.
(Which is why most big server deployments disable it.)
Doesn't have to be big, but without it, you can bring the system to its knees with 30% ostensibly free ram.
I don't find the term "performance" useful outside of a specific context. The main reason we don't use swap is predictability.
We know how much memory a particular VM should be using. If it exceeds that, failing quickly rather than changing behavior (slowing down) is far preferable, and then you correct whatever the problem is.
All the testing we've done with swap has shown it to have negative in server environments. I see it as useful for client machines, and maybe a bandaid to get by with underpowered systems if you have to for some reason. But that shouldn't happen in prod, especially if you're a public cloud user, as most folks here seem to be.
If your working set is bigger than your RAM you have a problem, swap or not.
However swap helps optimize your RAM usage so that your working set can fit with less waste. It allows the kernel more flexibility with what pages to evict from RAM. Without swap it can only evict pages backed by files. If you start evicting files that are in your working set you are just as screwed as if it starts evicting anonymous pages in your working set. With swap it can evict other unused pages before touching the pages in your working set.
I think there is some truth that it can be harder to notice with swap because there is more buffer between running great and literally crashing. However, in either situation you should be monitoring IO wait and application performance to ensure that you have enough memory for your working set.