I do wish people were more rigorous about changing swappiness. Maybe if swappiness=0 was documented as "pageout every byte of /usr/sbin/sshd and /lib/libc.so before you pageout a single byte of java's bloated heap" then people would be less eager to apply it.
Java's bloated heap, for example, has easy to use startup options that cause the Java process to use a deterministic amount of RAM, as does any other sane program that uses nontrivial amounts of memory.
OOMing (or swapping) a server almost always indicates an error that needs human intervention, so cleanly killing/rebooting the faulting system and raising the alarm in your monitoring system is the right thing to do.
Modern systems are usually teetering near the edge of using all physical RAM. This is by design; unused RAM is wasted. You want to use as much RAM as possible to avoid going to disk. What you call "OOMing" is when the applications on the system require more physical RAM than what exists. This is independent from the swappiness ratio.
The goal of virtual memory is to apportion physical memory to the things that need it most -- keep frequently used data in RAM, and pageout things which aren't frequently used. This makes things go faster. When you disable swap, you constrain VM's ability to do this -- now it must keep ALL anonymous memory in RAM, and it will pageout file-backed memory instead. Even if that file-backed memory is much hotter.
I came to this conclusion by following the kernel community, starting with the "why swap at all" flamewar on LKML. See this response from Nick Piggin http://marc.info/?t=108555368800003&r=1&w=2 who is a fairly prominent kernel developer. Nothing I've read from the horses mouth has refuted this since then. This is true even on systems with gobs of memory.
You're worried about systems which grind to a halt under memory pressure, which is unquestionably a concern. The thing is, disabling swap doesn't fix this. As soon as you're paging out important file-backed pages (like libc), your system is going to grind to a halt anyway. And disabling swap can't prevent this -- (same point from a VM developer here http://marc.info/?l=linux-kernel&m=108557438107853&w=2). To really give your prod systems a safety net, you need to (a) lock important memory (like, say, SSH and libc) in memory or (b) constrain processes which hog memory with (e.g.) memory cgroups. IMHO cgroups/containers are a better solution.
Ideal tuning would probably also reserve some decent amount of space for file caches and slab but I'm not aware of any setting that does that.
It's an entirely different story on laptops and development servers, where the workload varies widely, may contain large idle heap allocations worth swapping, and manually configuring memory usage isn't practical.
There is a school of thought which thinks especially server devs should just trust the operating system in this regard. One notable person of that school is phk of FreeBSD and Varnish fame: https://www.varnish-cache.org/trac/wiki/ArchitectNotes
But there are cases where it clearly does more harm then good. I had a postgresql datbase server with a lot of load on it. The server had loads of ram, more than what postgresql had been configured to use plus the actual database size on disk. Even so linux one day decided to swap out parts of the database's memory, i assume since that was very rarely used and it was decided that something else would be more useful to have in memory. When the time came for queries that used that part of the database, they had a huge latency compared to what was expected.
Maybe i'm misremembering and maybe there was some way of preventing that from happen while still having a swap enabled on the server..
Either way you handle it, the server is toast. So the real solution is to start dropping requests or downgrading service in some way to make sure the server never reaches OOM.
It definitely is important that the person signing the cheques knows how much capacity they're paying for, but that goes two ways. If you underprovision, they must understand that if they ever get mentioned in the NYT, their site will likely go down. It's up to them to balance the risks.
There's basically no reason not to have at least a small amount of swap. There's never been a benchmark showing no-swap as faster.
It should also be mentioned that things like mmap()+MAP_PRIVATE on a file will flat out fail if the file is larger than free_memory+swap. It's a common, easy and fast way to work with large files where the portion of the file you're working on gets paged in and out as required. Turn off swap and you break this functionality. You're basically limiting what you can do on your system with no performance gains.
Even with java apps (which I grant you are difficult to estimate memory limits on without knowing the application's design) swapping can be a useful way to page out unneeded parts of memory - and more importantly, keep needed parts of memory intact. This can also mean keeping the Java processes in memory while paging out apps which are less crucial, which leaves more room for Java, etc.
Swap is, for lack of a better comparison, the canvas sheet you use to catch someone jumping off a building, or a new york city sewer. In the first example you can use it to save your applications/servers so you don't need to reboot them (the higher your availability requirements, the less you can stand random reboots). In the second example it's the place you send those inhabitants you don't deem worthy of RSS.
The other thing you have to consider is the idea of memory overcommit in application and kernel design. The system is built with a promise that it has way more memory than it actually physically does. Applications will always reserve a fake huge chunk of memory and doesn't care that the system is lying to them about how much is really available. Without swap, when these apps attempt to use the nonexistent available memory, they crash. With swap, they survive.
Then there's apps designed to rely on swap like the Varnish malloc() and file storage methods, or database servers. Even if you disable swap there are still performance and stability problems related to the VM, and understanding this helps your apps run more efficiently. (http://blog.jcole.us/2012/04/16/a-brief-update-on-numa-and-m...)
That said, on my Linode I have swappiness=1 and a few GB of swap on the SSD since it is memory-constrained.
In any case, I've been running vm.swappiness=0 or 1 on nearly every machine I touch for many years now and have yet to see any problems. Swap is almost always a bad idea on modern machines.
Thanks for reading!
I currently have 1.2GB of swap in use on a machine that is not doing any active swapping at all; that's 1.2GB more space for caching.
[edit for spelling]
If you ran out of memory and malloc returned NULL that would be better, however a lot of applications rely on overcommit so there is no good answer:
* you can run with RAM + as much swap to make all applications happy and overcommit off, and accept the long delays caused by swapping
* run with just RAM, overcommit on and no swap, and accept that your applications may be killed by the OOM killer anytime
This is independent of applications whose working set size exceed that of physical memory.
If you only run applications you wrote yourself, you can prevent this behavior.
Except the downside of turning swap off is that you (possibly) don't cache your disk as aggressively, but the downside for having too much swap is that ill-behaved programs can grind your entire system to a halt for minutes/hours at a time by having too big of an active dataset. We're talking multiple seconds of latency, where you can't get anything done.
For me that's way too big of a downside for gaining a few extra MB of disk cache. I'd rather have it OOM right away.
96 GB would be plenty of memory for what I need to do--mostly I benefit from lots of filesystem caching. BUT--there's another user of this system and he does most of his work using Matlab. For the stuff he's doing, Matlab routinely sucks down 10s of gigabytes of memory. And it may stay that way for days at a time.
Without swapping, Matlab will happily sit on all that memory indefinitely, whether it's doing anything or not. Meanwhile, everything I do takes ages because of the paltry amount of memory left for FS caching.
With swapping, I can get some of that memory back for FS cache when Matlab isn't being used, and it makes a huge difference.
There can never be enough RAM. ;)
http://techreport.com/review/26523/the-ssd-endurance-experim...
I've an SSD that's been running in my development Linux laptop for very, very close to three years. The drive houses a couple of encrypted swap partitions along with the rest of the system. According to the SMART attributes, I've written 18TB to it in that time.
Don't worry about SSD wear. Really, don't. Either you'll get a drive that succumbs to crib death or super-shitty v1.0 firmware, or you'll get a drive that will last until long after you outgrow it.
People should of course learn about these setting as they change them, but many of the suggested values really are improvements for lots of use cases.