Almost Always Add Swap Space
haydenjames.io
haydenjames.io
Let me make it concrete: say 1 GiB of your application's RAM gets swapped out then is later paged in only as needed (4 KiB pages) with no readahead. Now the application's VM space is a minefield: there are up to 262,144 times it can stall for a ~10 ms disk seek (for a total of ~40 minutes). Sequentially reading in 1 GiB of RAM from disk would take only ~5 seconds.
Hopefully OSs use some readahead, but I'm sure Linux and macOS don't use enough. I find a system that has ever swapped to be totally unusable until I do "sudo swapoff -a" (on Linux) to force everything to be paged in, or just reboot (on macOS, I haven't found any other way).
Some of this can happen even without swap: the OS will still drop clean file-backed pages, so unless you mlock() your executables after startup (my production binary at work does this), you can still have major page fault-induced latency spikes.
Swap backed by SSD (or compressed RAM) should be more reasonable, but I've had enough bad experiences with swap on spinning disk that authoritative-sounding articles that encourage swap without mentioning this problem piss me off.
In contrast, if you don't have swap and run out of RAM, something will die and get restarted. In many cases this is a much better failure mode than continuing to run slowly.
I know I'm giving up some RAM that could be swapped out without impact but it's worth the stability of never swapping to an attached drive.
Windows would happily grow the pagefile beyond the limits I had set, and if the allocation loop was tight enough the OS chocked on handling all page faults. As in using literally minutes to respond to keystrokes.
By putting swap on a slow HDD, Windows has plenty of time to service mouse and keyboard input while waiting for the disk to respond, allowing me to kill the offending application before saturating my disk.
I assume this behavior was OK back when computers had very limited memory, and disks were slow anyway. But these days, if a single application wants more than the available memory (physical + pagefile), I'd rather the OS just said no to the application.
> But these days, if a single application wants more than the available memory (physical + pagefile), I'd rather the OS just said no to the application.
I agree.
You can adjust "swappiness" if it's moving stuff to disk too aggressively.
I disagree. The OOM killer usually kills the right thing, and if it doesn't a quick restart is usually better than continuing to run slowly.
You can also just put things in containers and set per-container limits.
> You can adjust "swappiness" if it's moving stuff to disk too aggressively.
What I want is to adjust un-swappiness: aggressiveness of paging stuff back in. Once the memory spike is over, it should page everything back in automatically.
Containers don't really solve your problem, it just adds a layer of indirection between the processes hogging RAM and the OOM killer, which in my experience adds even more uncertainty into what the OOM killer is going to do.
I won't say there's no circumstance in which the OOM killer is worse, but in general I disagree with you.
A stateful server application: maybe a standard ACID DBMS like PostgreSQL. I'd much rather it get killed (failing whatever transactions were in progress) and restart in a healthy state than continue to run slowly. Here I'm relying on a correct implementation of Durability, but if you don't have that, you have other problems.
Even more so for a quorum-based database server in which the leader dying (and another taking over) is much better than continuing to run slowly.
And yes, even more so in stateless servers, particularly when there are several of them. Going down and restarting quickly (a brief loss of capacity, a brief spike of errors if the client doesn't retry) is a lot better than continuing to be slow (a sustained loss of capacity, possibly leading to cascading overload; and user-visible latency problems if there's no hedging).
For something like a desktop application, it's debatable, but personally I'd still rather endure the up-front nuisance of having to restart it than have the OS try to paper over it until I finally realize why it's so painful and do something about it. (Insert anecdote about frog boiling here; I hear frogs actually do jump out before they boil, but yet somehow the story still rings true, if you know what I mean.)
Instead I would like to point out the ubiquity of things like nodejs, which make it all to easy for under-experienced developers to shoot themselves in the foot w/ a chaingun with respect to data design and management. It's not really the internal state of the nodejs process itself, it's the external state that it is inevitably being manipulated in the node process.
Some times that luxury doesn't exist though and you just have to get shit done and hope playing fast and loose doesn't come back to bite you.
Secondly, the thing with swappiness is that it alters the balance between stuff that is cache and stuff that is live data. But "stuff that is cache" includes your running software text, libraries, essential operating system bits, etc. If you set swappiness too low, then in a low memory situation the OS will throw away all the code that it is trying to run instead of swapping out some of the live data. That is what many people don't realise.
And using swap as some kind of "RAM emergency cover" does not make any sense to me. Personally, I have never had a case where I could gracefully recover a host that had started swapping.
First, my production servers are not doing many other things besides being production servers. It's not like they are running a bunch of unnecessary services, and if I kill some then my application can recover. If a process is out of control then it's almost assuredly something important that I'm going to have to kill anyway.
Second, I find it's much harder to detect degraded performance than it is to detect a dead process. It's very, very easy to have a health check that will detect a host who has stopped listening and drop that host from the LB. And alerting on that scenario is very easy as well. The alternative is a host that's operating in a degraded state, which I need to detect with more sophisticated health check + alerting, and in the end my resolution is just going to be to kill everything anyway.
In a properly designed HA environment the loss of a host should be no big deal. Architecture should be focused on making sure a host goes down ASAP if it's having problems, not letting it survive in some kind of zombie state.
The memory freed from a page that got pushed to swap might be better used for serving disk cache, for example.
If you don't have enough RAM (even for spikes), a preemptive push when things aren't busy might mean that when they do become busy, you won't have to wait.
At least, that's the theory as I understand it. Whether or not it works in practice, IDK.
> And using swap as some kind of "RAM emergency cover" does not make any sense to me. Personally, I have never had a case where I could gracefully recover a host that had started swapping.
I have successfully recovered hosts that began swapping, and even hosts that filled their swap. But it's beyond painful, and in most of the cases where I have a host doing such swapping, an OOM kill would have been a welcome reprieve.
> Second, I find it's much harder to detect degraded performance than it is to detect a dead process.
Agreed, and restarting something that's been OOM'd is simple enough too. Our hosts tend to stop reporting metrics when they start swapping, simply b/c so little is actually getting done, so they get labelled as "down". Linux also exposes a metric called "major page faults" (IIRC, it's "pgmajfault" in /proc/vmstat) that records the number of major (required service from disk) page faults; if the rate of that value is too steep for too long, that ≅ swap thrashing.
For this reason, you want an OOM killer to kill stuff way earlier than when you actually run out of RAM, because if you ever reach that point it is too late. I use a program called earlyoom, which has saved my server from having to have the big red button pushed quite a few times, when fumble-fingered PhD students "accidentally" consume all the RAM in their pet projects.
Along those lines, take a look at this: [1]
> Pressure stall information were added to Linux 4.20 as a way to quantify resource pressure in the system in a better way than the traditional load average. PSI aggregates and reports the overall wallclock time in which the tasks in a system (or cgroup) wait for cpu, io or memory.
> This release [Linux 5.2] lets users to configure sensitive thresholds and use poll() and friends to be notified when a certain pressure thresold is breached within a user-defined time window. With this mechanism, Android can monitor for, and ward off, mounting memory shortages before they cause problems for the user. For example, using memory stall, monitors in userspace like the low memory killer daemon (lmkd) can detect mounting pressure and kill less important processes before device becomes visibly sluggish. In memory stress testing psi memory monitors produce roughly 10x less false positives compared to vmpressure.
I haven't tried it yet, but it sounds promising.
[1] https://kernelnewbies.org/Linux_5.2#Improved_Presure_Stall_I...
> If you don't have enough RAM (even for spikes), a preemptive push when things aren't busy might mean that when the do become busy, you won't have to wait.
> At least, that's the theory as I understand it. Whether or not it works in practice, IDK.
These are the ideas that I hoped the original post was going to be about.
Disk cache is a good one. I'm not knowledgeable enough about this, will the Linux kernel automatically start using RAM as disk cache if there's enough free space? Is it smart enough to prioritize disk cache over infrequently used pages and then know to discard the disk cache if it does need to page stuff back in from swap?
> will the Linux kernel automatically start using RAM as disk cache if there's enough free space?
Yes, exactly.
> Is it smart enough to prioritize disk cache over infrequently used pages and then know to discard the disk cache if it does need to page stuff back in from swap?
It is, and this is what's controlled by the vm.vfs_cache_pressure sysctl that the post discusses how best to configure.
Putting rarely used (or never to be used again) pages in swap frees up physical RAM to be used as disk cache, which can improve performance for file access. If you have enough RAM to cover all application memory, and cache your entire filesystem, there is no gain to be had.
I agree that slowing down low memory scenarios is silly though. In many cases killing a process is easier to deal with than unexpected poor performance. The slow death of out of memory systems doesn't seem to be unique to having a swapfile though; Linux seems to discard any recoverable pages (such as application / library code backed by files on disc) in the run-up to OOM, resulting in thrashing.
Ideally I'd want to have swap, but automatically kill processes when performance deteriorates due to swapping, rather than when the system is strictly out of memory.
The only way that makes sense is if there is some cache the kernel has which is faster than RAM. BUT the kernel cannot move things between this cache and RAM, it can only treat the cache + RAM as a single blob, so to free up space in the cache the kernel has to move pages to disk. So by having some swap space the kernel can move hardly ever used pages to disk which is the only way to free up space in this cache.
It's possible that what I just described really exists, I don't know that much about the internals of how the kernel manages memory. But if that is what's happening then THAT is the detail that would have been useful in this post.
I had a small, very lightly loaded VPS that had no swap. Occasionally, it would not start all its services on boot, in particular the db, due to OOM.
Adding a small amount of swap allowed it to boot cleanly. The swap was not subsequently used.
The alternative was to upgrade the server RAM at double the monthly cost.
One thing I think the author was getting at, but never explicitly stated is that we spec most things for the average expected workload with a bit of headroom--not the peak workload since that is very difficult to assess. Peak workload might also be few and far between and not worth paying for upfront. Swap allows you to deal with peak workloads more gracefully for many applications.
Having your process randomly die isn't a big deal with stateless processes (which is a more and more common design), but really sucks for stateful things. Not just the possibility of lost/corrupt state, but process startup and loading of state can take a long time.
We have a few servers that don't have swap and I was hoping to find an excuse to add it in this article. Instead, it reenforced that decision because of type of things we run on it.
This is almost always false in practice though. Peak workloads means more than just lots of memory, it means high CPU and high IO (disk and network). If you under-spec your servers such that they need to hit swap during peak loads they are going to generally cause more problems than if they just die. A server dying is something you have to prepare for and deal with anyways, and is much simpler to plan for than the weird latency issues you'll get with a server hits swap.
Think of it as similar to Erlang's fail-fast model. Better for things to just die and be recovered than to continue on in a bad state.
If your hardware is underspeced for the workload, there's still a business case to be made before you buy new hardware. Sometimes the preferred failure mode is to slow down, other times it's preferred to die and restart. Hardware speced for peak loads can sometimes cost multiple times otherwise appropriate hardware...assuming you even know what peak loads will be.
I have one class of application that is fairly idempotent, we don't have hard drives on these machines, so we don't use swap and expect it to die if it exceeds the amount of memory requested. This article reenforced this.
I have another class of application that queues up things to process. The specs work on the current hardware 99% of the time, but on holiday weekends when large projects are delivering things slow down because work is submitted and people leave so nobody cares. Reloading the state from disk can take 20m+ and without swap if you were to exceed memory you're dead in the water until you purchase more hardware. Our contingency hardware has identical specs as the original.
But that still relies on uncorrelated failures.
Otherwise that seems like an orthogonal discussion to swap vs. OOM killer. It's something to be aware of and you need circuit breakers in place but I don't think it's an argument for letting hosts swap instead of having the OOM killer deal with it.
Using swap space may make sense for a traditional physical, stand-alone server, where having it fail quickly doesn't really help. Or maybe for a workstation.
Why should I care if I always have enough RAM? Gedankenexperiment: Let's say u'd have infinite RAM, would swap make any sense?
[1] You can add in such a limit via RLIMIT_AS, but it'd have a lot of collateral damage like preventing LMDB from mmap()ing a big database. I don't recommend this.
[2] I'd recommend setting limits on a cgroup; with a systemd unit file, see MemoryMax= and friends in the systemd.resource-control(5) manpage.
On Linux you can make them fail, but don't do that unless you know what you're doing.
The theoretical worst case scenario is that the system hits an OOM situation and you end up with data/filesystem corruption... I have not had this occur due to swap being turned off this century.[1] The typical worst case scenario is that either due to a runaway process or accidentally overloading the system for a indefinite period of time, the system goes completely unresponsive so you have to perform a hard reboot. This happens to me 1-2 times a year but I have yet to loose any data (and if it ever does, that's one of the many reasons to have backups) The more common scenario is that a runaway process will OOM briefly, when it can't allocate the memory it needs crash, and everything is back to normal within a couple of seconds.[2]
FWIW, mobile (Android and iOS) devices typically run without swap just fine. IMO, this 'always enable swap' is bad advice and dates from a time when very bad things could and did happen if you OOM'd a system. The marginal performance improvement from freeing up on the order of megabytes (or even tens of megabytes) of RAM on a system with double digit GB's of RAM makes absolutely no sense to me. It makes even less sense on SSD systems where you would be unlikely to notice a chronic swapping situation until you had severely shortened the life of your SSD.
[1] This would be a good reason to not turn off swap on a production database server where it is more likely to be an issue, for example.
[2] One of the dubious advantages to 'modern' programs not handling errors like OOM is that, since they don't check for it, they can't handle when it happens and just crash. This was actually more of a problem in the 90's and earlier when it was more common for applications to handle OOM situations which could actually make the problem worse since they would often hang around waiting for any memory to be freed and use it which resulted in a persistent OOM situation.
I'd suggest trying it and seeing what happens, if you wish to learn rather than listen to people who aren't using your host - and don't know what you run..
In a Microsoft OS? Depends on the app -- some are dependent (many games) on having swap available (and may err silently or under false pretenses without one).
I've found that a lot of MMO style games include cheat detection systems (see : rootkit) that rely on swap, and fail without it. As to why they do that, i'm unsure.
If you have swap, this happens later and more gradually, and lasts longer before the OOM killer kicks in.
For example on my current laptop if I hibernate or sleep the system the USB fails to initialize in Windows until I power-cycle. I think this is mainly due to me having about 10 or 15 USB devices plugged in via hubs and USB ports. But that's not the norm.
I've also got 64 GB of ram and a pretty decent pair of NVMe SSDs so read speed isn't really a limiting factor for startup time.
https://github.com/armbian/build/blob/master/packages/bsp/co...
zramctl output from a tiny NanoPiNeo2 with 1GB Ram, running from a 64GB Sandisk SDXC card:
NAME ALGORITHM DISKSIZE DATA COMPR TOTAL STREAMS MOUNTPOINT
/dev/zram1 lzo 496.7M 4K 76B 12K 4 [SWAP]
/dev/zram0 zstd 50M 2.5M 449K 832K 4 var/log
Edit: Formatting (useless)
If memory pressure is low, and the swap is used to page out rarely / never used pages in favour of more disk cache, then there is a performance gain to be had, and wear should be minimal.
You can encrypt swap with a randomly generated key to aid security.
Ideally memory compression and disk swap would be used together; there's a gain to be had by shrinking lesser used pages, but no sense in taking up any space in RAM for pages showing no sign of life.
In my experience, when Chrome is being a memory hog I can swap about 2/3 of its memory out with zero performance impact. As long as I can swap that away, I never run out of memory. If I can't swap it to disk, I have problems. I don't think that qualifies as "too little ram".
So in this case performance is fine, security should be an issue of encryption not disabling, and the SSD is not going to wear out because of an extra <1% average utilization.
I'm prefer add RAM as much my tasks needs it.
What you really want is to never have a severe out of memory situation, and you can do that either by getting enough RAM that you never use it all, or by killing stuff nice and early when memory starts getting low.
The server never be unresponsive 'cause lack of swap. Don't like the logic "we just add a swap and maybe it will be slower a bit, but still working".