Do we really need swap on modern systems?
redhat.com
redhat.com
* If you need a low-latency server or workstation and all of your processes are killable (i.e. they can be easily/automatically restarted without data loss): disable swap.
* If you need a low-latency server or workstation and some of your processes are not killable (e.g. databases): enable swap and set vm.swappiness to 0.
* SSD-backed desktops and other servers and workstations: enable swap and set vm.swappiness to 1 (for NAND flash longevity).
* Disk-backed desktops and other servers and workstations: accept the system/distro defaults, typically swap enabled with vm.swappiness set to 60. You can and likely should lower vm.swappiness to 10 or so if you have a ton of RAM relative to your workload.
* If your server or workstation has a mix of killable and non-killable processes, use oom_score_adj to protect the non-killable processes.
* Monitor systems for swap (page-out) activity.
* vm.swappiness = 0 The kernel will swap only to avoid an out of memory condition, when free memory will be below vm.min_free_kbytes limit.
* vm.swappiness = 1 Minimum amount of swapping without disabling it entirely.
* vm.swappiness = 60 The default value.
* vm.swappiness = 100 The kernel will swap aggressively.
This is not the case.
It used to be the case, but changed in kernel version 3.5-rc1 (2012 ish)
There was a discussion about this on HN a few weeks ago: https://news.ycombinator.com/item?id=13511086
And there's a blog post on the percona website about how this rather bizarre change bit them: https://www.percona.com/blog/2014/04/28/oom-relation-vm-swap...
I call it bizarre because (as I wrote in that other HN thread) a) it changed the behaviour of lots of production systems in a surprising way, and b) if you want to ensure your processes never swap you already had the option to not have a swap file or partition.
(Maybe it wasn't so new... https://www.percona.com/blog/2014/04/28/oom-relation-vm-swap...)
mysql_enable="YES" mysql_oomprotect="YES"
Now every time you start the MySQL service it's automatically protected
oom-kill-protect fromenv
And the conversion from something like OOMScoreAdjust is quite straightforward. A PostgreSQL systemd unit file that read OOMScoreAdjust=-625 becomes a run program that contains oom-kill-protect -- -625
* http://marc.info/?l=freebsd-hackers&m=145425153624976&w=2* http://jdebp.eu./Softwares/nosh/guide/oom-kill-protect.html
There is also zram (just swap in memory lz4/lzo compressed) and zswap (compressed cache in memory for swap pages before hitting disk) that needs a real swap device but compresses pages beforehand.
I run zswap on my Desktop and on a few servers and it gives you some more time before the oom killer comes and the system feels a bit longer responsive.
zram is a nice idea but quite a beast in practice (at least on MIPS with 32mb RAM) sys constantly at 100% if you need it and other quirks. Maybe it got better or I did something wrong.
But if you need an in-memory compressed block-device it's pretty great - you can just format it with ext4 and have a lz4 compressed tmpfs.
I set it up with one zram device per CPU core for a total space of ~20% available RAM.
No performance issues w/ zram so far so I haven't felt the need to change the compression algorithm.
# modprobe zram num_devices=1
# echo 1G > /sys/block/zram0/disksize
# mkswap zram0 /dev/zram0 -L zram0
# swapon -p 100 /dev/zram0
Official documentation here: https://www.kernel.org/doc/Documentation/blockdev/zram.txtUntil you actually run out of memory, zram seems very much a set-and-forget type of thing. No babysitting required.
tl;dr: it does what it says on the tin, and ... with minimal cpu impact.
Zswap maintains default kernel memory allocation behaviour, with the tradeoff that it needs a backing swap device to push old pages out to (which is why zram tends to be used more often in embedded devices that only have a single volatile memory store, of devices with limited non-volatile storage).
Is this that big of a worry? I have a 5-year old SSD in my daily driver laptop, on OS X which loooves to swap out anything it can to gain memory for disk cache, and I'm still barely 15% into the SSD wearout.
What seems to keep swap alive is that asking for more memory ("malloc") is a request that can't be refused. Very few application programs handle an out of memory condition well. Many modern languages don't handle it at all. Nor is it customary to check for a "memory tight" condition and have programs restrain themselves, perhaps by starting fewer tasks in parallel, opening fewer connections, keeping fewer browser tabs in memory, or something similar.
I've used QNX, the real time OS, as a desktop system. It doesn't swap. This make for very consistent performance. Real-time programs are usually written to be aware of their memory limits.
Most mobile devices don't swap. So, in that sense, swapping is on the way out.
These aren't mutually exclusive and are actually complementary with swap.
If you have more than enough memory then swap is unused and therefore harmless. The question is, what do you do when you run out? Making the system run slower is almost always better than killing processes at random.
And it gives processes more time to react to a low memory notification before low turns into none and the killing begins, because it's fine for "low memory" to mean low physical memory rather than low virtual memory.
It also does the same thing for the user. "Hmm, my system is running slow, maybe I should close some of these 917 browser tabs" is clearly better than having the OS kill the browser and then kill it again if you try to restore the previous session.
..which operating system is that?
- notice system starts swapping (if you do not monitor that, to me it sounds as careless as driving on the highway on 2nd gear and ignoring engine noise -- ideally the OS could proactively help here but unfortunately I don't know a good "automated" tool)
- find out which process/app uses the most memory (Linux can even tell you which ones use the most swap space [1])
- decide which one you want to (gently|forcefully)-(quit|restart|whatever). Exercise judgment.
[1] http://stackoverflow.com/questions/479953/how-to-find-out-wh...
In practice, heavy swapping (forth and back) makes it impossible to even kill the culprit manually (because I can't open an xterm or whatever). While there is often no benefit to have the processes continue running that slow.
Also, idealistically programs should be written with the assumption that the machine could go down at any instant. Having a few more cases where the program is killed will have the effect that the program is better tested and debugged.
I'm not sure that turning this into a market-analagous operation (bidding some ... other scarce resource -- say, killability?) might make the situation better or worse. And the problem ultimately resides with developers. But as a thought experiment this might be an interesting place to go.
Niceness allows for higher-priority processes to preempt others, but doesn't address the problem of an overwhelmed queue.
And processor scheduling isn't memory allocation. Time is ultimately some percentage of wall-clock (and/or overcommittment). Memory is ... different.
There's also the question of such stuff as garbage collection and scheduling of that. I had the opportunity to do some JVM tuning "ergonomics" (horrible name) a few years back. Turns out that you get far better behaviour in most cases by decreasing the sweep frequency and increasing the the allocation chunks (terminology is escaping me), due to the fact that natural attrition deallocates memory, and running sweeps too frequently simply chews up massive amounts of CPU time with no return on freed memory.
We also identified processes which genuinely did require very large memory allocations, and allocated hardware specific to those.
Specific workflow and process understanding (always idiosyncratic to a particular work assignment) was necessary, and took time to acquire.
(For example, if a process forks, a great deal of memory is shared between the parent and child. In theory one process could dirty all of their writeable pages, forcing the kernel to allocate a second copy of each page. In practice, almost no process that forks will do that and reserving RAM (or swap) for that eventuality would require you to run significantly oversized systems.)
My hunch is that the OS is swapping stuff back in stupidly. Once memory is available, I'd like it to page everything back proactively, preferring stuff from swap and then from file-backed mmaps. But instead it seems to be purely reactive, each major page fault requiring a disk seek to page in what's needed with little if any readahead. Basically the whole VM space remains a minefield until you stumble over and detonate each mine in your normal operation. Much better to reboot and have a usable system again.
On my Linux systems, I've turned off swap.
On OS X...last I checked, I wasn't able to find a way to do this. I'd like to turn off swap entirely, or failing that, have some equivalent way to force all of swap to be paged in now so I don't have to reboot when I hit swap. Anyone know of a way?
That depends. If your workload exceeds the amount of available memory, you will start "thrashing" the disk and that can make a system un-responsive.
If you happen to launch a large application, or start working with a big file, unused pages will be evicted to disk to make room and, after some slowdown, the system should become perfectly usable again. YMMV
On OSX, I don't know a way, but I can't recall the last time I had to reboot due to RAM/swap issues, even when I was developing apps on a 4GB Macbook Air. I guess memory compression, which is enabled by default, helps here. Most OSX systems have very fast SSDs as well.
http://osxdaily.com/2015/10/29/use-trimforce-trim-ssd-mac-os...
I had a client whose app generate PDFs that would cause this to happen.
I'll always remember when I used to load a large piano instrument in a VST DAW on Windows 7, taking about 3-4GB out of 12GB of RAM, it played perfectly fine but if I left the application open, invariably on the next day I'd get a barrage of audio dropouts when pressing any new piano key. One trick was to put my arms on the entire keyboard a few times to wake the swap back up to memory. Another trick, which I ended up relying upon despite occasional low memory warnings, was to disable the pagefile entirely - that sure fixed the problem.
I'm not sure how/if things have improved since with Windows 10 and SSDs, but I always felt there was something wrong with the algorithms, since even with GBs of memory free at all times, old memory content would tend to end up on disk, without any good reason I could see.
I assume the OS used time to prioritize various caching/pre-fetching techniques over actual application data, and/or once paged, never preemptively loaded data back to RAM even if plenty of memory was available.
What is an unused page? One that the foreground, memory-hungry application doesn't need? Okay, fine, but what happens when you switch back to some other application? My experience is that it needs the RAM that was paged out, and it doesn't get paged back in all at once. Every time you hit some 4 KiB of memory that happens to be paged out, you wait another 10 ms. I don't know how much beyond the 4 KiB gets paged in at the same time. Worst-case, there's no read-ahead at all. Let's say the application is using 1 GiB of RAM. Then this can happen 262,144 times, which means 44 minutes of waiting in small bursts as you're trying to use it, rather than the 10 seconds (at 100 MB/s) it'd take to read it all in one go. That's what I mean when I say the machine is unusable.
The other culprit I suspect is too many potentially blocking processes Cinnamon's main thread, but I've never figured out a decent way to go after the problem.
The OS should start swapping very early to avoid bursts of disk I/O and rendering the system unusable. On linux this is somewhat configurable, even if not user friendly, but a combination of swappiness and vfs_cache_pressure could turn it into a usable machine, taking care of inefficient memory usage, memory leaks, unnecessary vfs cache, etc.
I think you're talking about the I/O of paging dirty things out, but I'm talking about the fact that some memory location is no longer present in RAM, so accessing it will take 10 ms or more to page in.
The system is not only useless while actively swapping. It's useless after it has ever swapped, and you can only recover by disabling swap ("sudo swapoff -a") or rebooting.
sudo launchctl unload -w /System/Library/LaunchDaemons/com.apple.dynamic_pager.plist
Then:
rm /var/vm/swapfile*
Disclaimer: haven't tried doing it this way on Sierra.
20 years ago on Windows 98 it just started swapping, but it was no big deal. If something became too slow to be usable, you could just press ctrl+alt+del and kill that swapped program and everything worked fine afterwards.
On the other hand, my modern linux laptop, it starts swapping, and it swaps and swaps and you can do nothing, not even move the mouse, till 30 minutes later something crashes.
EDIT: see comment below for more accurate numbers.
>A typical reference to RAM is in the area of 100ns, accessing data on a SSD 150μs (so 1500 times of the RAM) and accessing data on a rotating disk 10ms (so 100.000 times the RAM.
Latency Numbers Every Programmer Should Know
Default settings for dirty ratio and dirty background ratio exacerbate the issue: more data is held onto before it is written, and once the background ratio is hit, any application writing to disk will block.
I feel like Linux has, in general, from a UX point of view, the worst behaviour when swapping and the worst behaviour in general under memory pressure.
I feel like it has gotten worse over time, which might not be just the kernel but the general desktop ecosystem. If you require much more memory to move the mouse or show the task manager equivalent, then the system will be much less responsive when it thrashes itself.
Honestly, I'ld much rather have Linux just crash and reboot, that'd be faster than it's thrashing-tantrums.
Luckily, there's earlyoom, which just rampages the town quickly if memory pressure approaches. Like a reboot (ie. damage was done), just faster.
In any case, it makes me sad (in a bad way) to see how bad the state of things is when it comes to the basics of computing, like managing memory.
i3 ≥ v4.13 will pick up this value from the Xft.dpi resource in ~/.Xresources as well, which is the more common way of configuring DPI.
edit: haven’t tested this within Xwayland, though. Note that i3 is only supported on X11.
Yep. I bought a second computer for full browsers. One for dev, another for 'full' browsing (Javascript on) and on my i3 dev machine, I only have NoScript browsing on for dev stuff.
sudo swapoff -a sudo swapon -a
At that time, swapping out a 4k page was a significant part of memory: 4k of 16 MiB is 1/4096 of memory. Each swap gets back a lot of memory the program needs. Now the swap is still 4k pages, but memory has expanded by a thousand fold. Basically swap is a thousand times worse today than it was in the time of Windows 98.
For harddrives swap isn't used now to expand memory, it's used to remove initialization code and other 'dead' memory. Swap should be set to only a tiny fraction of the memory size for this reason, to prevent it from being used to handle actual out-of-memory conditions. But realistically for most users it's not even worth enabling at all because of the occasional memory that needs to be swapped in from disk.
For SSDs the seek speed has improved to match the extra memory so swap can still be used like in the old days to expand the effective memory size. But memory is so large a swap file that's a fraction of memory size to offload 'dead' memory is enough unless there's a specific reason to actually use swap for out-of-memory.
I suppose it might be possible to whip something up with cgroups and policy that will keep the VT, bash, X and a few select programs always resident in memory and give them ultimate I/O priority, but I haven't tried.
https://kernelnewbies.org/Linux_4.10#head-f6ecae920c0660b7f4...
Personally, I've never given VMs swap. I'd rather have memory pressure trigger horizontal scaling (or perhaps vertical rescaling, for things like DBMS nodes) than let Individual VMs struggle along under overloaded+degraded conditions.
There can be no consensus because there is no one answer.
Not sure what you're referring to here. This story doesn't recommend eliminating swap...
So, it doesn't exclusively recommend it, but it concedes that there are use cases where it makes sense.
On Windows without swap when you hit a remotely low on RAM point, things start going really poorly for some reason - random latency. So with 16 GB of RAM even I can't disable swap on Windows without some really strange performance characteristics, I run SSDs so I really wanted it off and I just stuffed more RAM in my box - with 32 GB it isn't a problem.
On Linux however, you can pretty much turn it off and everything will run smooth until you're actually out and then you lag badly briefly, Linux's oom-killer does its thing and all is good again within the span of a few seconds.
If you ASK about swapping on windows, you get people telling you that "Microsoft engineers are smart, don't disable swap and go <insert expletive here>" even if you asked something that is NOT about disabling swap.
So, I had this gamer laptop, i7, nVidia GPU, 8GB of RAM (when most machines had 2 or 4), but some stupidly slow 5k RPM HDD made for power saving and locked "noiseless mode", thus very slow seek too (ie: it moves the heads slowly to avoid making noise and for aerodynamic reasons).
I noticed that ever after I just booted up, RAM usage would jump to 6gb and the HDD would trash endlessy and make the machine unusable... after some research I found some interesting posts by MS employees about it:
Windows can "preemptively" use swap, it will write on swap things it thinks you might need to swap out. Sounds good on paper.
Also, Windows has several caching systems, that will write to "RAM" random crap.
One day that was particularly bad, I noticed that when I booted, Windows would immediately attempt to copy to RAM a gigantic binary file that was the sound files of a game I played a lot recently, this caused trashing due to reading the file, then, it would attempt to load other programs it had to, then page out immediately, and enter some crazy loop of trashing the I/O forever... Every time I opened the task manager and looked at the graphs, disk I/O would be constantly maxed out at 100%...
Disabling the VM made the laptop behave better (despite all the bugs Windows have when you disable VM).
But what I really wanted, was to change how the VM works... I wanted to keep the VM, and the caching, but change settings, for example I would set it to NOT page out anything at all unless RAM was used more than 80%, and also to never "cache" stuff unless HDD was actually idle and a good amount of RAM free. But sadly, this can't be done it seems, I got no useful answer on stackexchange sites when I asked this (But got a couple personal messages and e-mails full of expletives in many places where I asked about it, for some reason people get personally offended when the subject is virtual memory).
Oh, that's just superfetch, it's a service you can disable to reduce a bit the idle trashing after desktop has loaded.
Because of this memory pressure will be higher on a windows box. Pageing helps paper over this as the commit can be billed to the page file not RAM. Windows is smart enough to not write anything to swap until you actually use the page so in practice this is rarely a problem.
The benefit to this approach is you actually have a hope of recovering from OOM.
That's not true, you have to turn vm.overcommit_memory on in Linux for that to happen I believe. Which is off by default in most distros.
The default is to allow "sensible" overcommit whatever that means. From my experience whatever "sensible" is, really is sensible and I haven't had issue with that. You can also set it to allow all memory allocations, even "silly" ones (i.e. allocate 100GB memory on a system with 10GB RAM); or to refuse overcommiting memory.
Usually selecting sshd to kill, in my experience, rendering the server inaccessable.
* On a laptop to hibernate, which results in zero power consumption vs suspend which will drain the battery in a day or so
* I use tmpfs for /tmp and using swap as the backing is far more performant than regular filesystems
This seems absurd. You're running an in-memory filesystem backed by memory-on-disk? You weren't comparing to a journalled filesystem or something like that?
I consider files in /tmp to be temporary, and do not expect them to survive a reboot. (Actually I prefer they don't - less administration and housekeeping.) They also have random lifetimes ranging from fractions of a second to several days. And random sizes from zero length to gigabytes (eg making an ISO image).
With tmpfs RAM is used which provides the best performance since the filesystem is trivial. Memory pressure will cause swap to be used as needed. Files not accessed will end up in the swap and taking no RAM.
By far the fastest I/O is the I/O you don't have to perform.
If you're using a swap partition just for the sake of /tmp then it's the same difference, no?
No. The big difference is that regular filesystems try to do I/O to their backing device - heck that is their point, and what they do the vast majority of the time. tmpfs does not do any I/O. However I/O will happen when there is memory pressure by the swapper, but that is going to be rarer.
ie with tmpfs, swap is a spillover mechanism. With a regular filesystem, the underlying device is the primary mechanism.
Swap can also be used for actual swap on the occasions it is helpful.
Yes and no - aren't they just two ways of looking at the same decision? Regular filesystems will buffer, and when the system is low on memory it will flush buffers, using similar criteria to deciding whether to swap.
> Swap can also be used for actual swap on the occasions it is helpful.
Swap enables actual swap, sure. My experience is that it usually hurts more than it helps though.
That is the bit you are missing. Unwritten filesystem data is regularly flushed. The flush interval is often around 5 seconds. Lookup "pdflush" to get the gist, although things have changed since then. Same with laptop mode.
Quite simply if a file is created and lives for at least N seconds then there will be disk activity irrespective of memory pressure. N is 5, perhaps up to 30 seconds in normal use.
Even if the file contents aren't fully flushed, metadata is.
Swap is not strictly needed for this:
(it boils down to vm.swappiness=1)
https://wiki.debian.org/Hibernation/Hibernate_Without_Swap_P...
Still, it's possible to hibernate without having the drawbacks people have been complaining about in this thread.
for sure it is. Even for swap files inside an encrypted volume.
https://vadim-kirilchuk-linux.blogspot.com/2013/05/swap-file...
The crux is knowing the swap-file offset and passing that argument as resume_offset on boot.
Have a look. It's totally doable and you won't need to have an actual partition.
I sort of wonder if we'll see a 100% RAM, large memory laptop soon that boots from an SD-card or in a cryptographically secure fashion over 4G wireless networks, aggressively disables RAM for power saving and suspends well.
Once a server hits swap, it's dead. There is no recovering it other than for exceptional cases. If you are swapping out, you've already lost the battle.
I tend to configure servers with 512MB to 1GB swap simply so the kernel can swap out a couple hundred MB of pages it never uses - but that's really more to make people feel better than it really being useful at all.
(The limitation came about because the simple way to handle swapping is to assign every potentially swappable page of virtual memory a swap address when you allocate it in the kernel. Then the kernel always knows that there's space for the page if it ever needs to swap it out and you're never faced with a situation where you need to swap out a page but there's no swap space left.)
If RAM and DISK are the same, then writing a file system is just writing an in-memory tree. No need to pull data from the disk, just navigate the tree in your program's memory and pull the blob data out. Want to persist acorss reboots, protect against power outages, or save user settings? Just set a variable and it'll be there.
The benifits are much better then the costs.
[0] - https://web.archive.org/web/20031029002231/http://www.eros-o...
https://en.wikipedia.org/wiki/MUMPS
Setting data in memory is the same as setting data on disk, the only difference is the name of the variable:
s X=1 ; store 1 in variable named X, in memory.
s ^X=X ; store 1 in variable named X, on disk.
s X=^X ; load disk to memory
Note that EROS is not providing a write-through cache. It's providing a write-back cache using checkpointing coupled with a journalling capability and ability to explicitly sync data.
So it's leaky: Your application needs to know that it needs to structure it's writes to memory so that they will make sense if the system comes back up with some of the data missing, and needs to know how to use the journalling functionality.
It can't just act as if it's running in RAM forever.
If you just pretend you're running in RAM, and the system crashes, you will lose the data between the crash point and the last checkpoint unless you have explicitly used the journalling mechanism. Often that is acceptable. E.g. since you're restoring the program at the same point in time, if the changes are entirely based on data that were in the system at the point of the last checkpoint, it will just redo the work to calculate the changes.
But if there are side-effects, that is often not going to be acceptable. E.g. database updates that the system has said were committed will suddenly disappear.
To solve that, EROS has a journalling mechanism to allow you to give guarantees about specific data that changes in between checkpoints, but that requires applications to explicitly use it to tell the OS what needs to be saved when so that the application can guarantee that a given piece of data has been durably recorded when it promises a client it has been recorded, and that the writes get correctly ordered.
That's a sensible compromise - if you do it right, it only needs to touch the "boundaries" where the system does IO.
> But if there are side-effects, that is often not going to be acceptable. E.g. database updates that the system has said were committed will suddenly disappear.
Yes, not every system will forever be recoverable. You can definetly crash at just the right moment to ruin your year. I'd still like the other safety constraints that this provides because I still think that even if N (where N is the number of threads available on the CPU) processes have a chance of being corrupted we're still going to be saving the rest of the processes on the system that aren't currently transacting with one another.
The point is more that it doesn't mean you can just treat things like we do RAM now. It ends up being closer to how you'd work in apps that use mmap'd files to back persistent data.
The capabilities model was interesting too.
I think my biggest problem with a persistence model like that, though, is that I suspect it would encourage not thinking seriously about state, and ideally I'd prefer a system where state is minimized. E.g. compare EROS with the virtual opposite: Android, for example, where apps might find themselves killed anytime. Some apps handle it poorly and go through lengthy initialisation processes again when restarted, but many maintain the illusion of being fully presistent to the extent that users can rarely tell if they've started from scratch or not.
I'd love to see more work into allowing the illusion of a persistent process with little developer effort, though. Perhaps OS-level per application checkpointing support that the application can have control over (allowing the app to control exactly what gets checkpointed, and when to ditch the checkpointed state to reinitialise). So "cleanup" can occur by restarting processes transparently for users while hopefully providing most of the benefits of persisting state.
Perhaps coupled with OS-level checkpointing of the information required to bring said "virtually-persistent" apps back up in the same state after a reboot.
Frank Soltis' book is recommended reading: https://www.amazon.com/dp/1882419669/
It's a shame hardly anyone knows about it. Those things are a joy to use. You can get a free (limited, but still useful) AS/400 user account to play around with at http://pub400.com/. I really recommend it.
(Disclaimer: I'm slowly working on a system that resembles AS/400 in many ways, but optimized for analyzing and reporting on very large timeseries databases. It's intended for business applications that require a combination of scheduled reports and fast ad hoc analysis of big timeseries data, initially the oil & gas industry (which is where I work in my “real job”).)
These are all problems that have been solved on EROS based system. They used to do demos where they would setup a system and have someone start working on some code or a text document, they'd pull out the power plug of the system, plug it back in and the user would be right were they left off. No data loss, no corruption, just back to work.
None of that was handled in user space. That was all opaque and you didn't have to worry about it at all.
I'm guessing the system must effectively "checkpoint" your work regularly and sync to disk to avoid partially saving a state and corrupting the data.
This isn't terribly different from working on a memory mapped file except that it also saves the ephemeral state of the running program so it can be restored. But I still don't understand how it's not going to be horribly slow when you start your program and the first thing it does is allocate 4GB of memory for its workspace. Synching all of that data to disk is a massive undertaking, and this isn't an uncommon use case, people start virtual machines all of the time.
And paging is slow as balls. That's what this whole article is about, modern machines are unusable when they start paging.
> Although OS X supports a backing store, iOS does not. In iPhone applications, read-only data that is already on the disk (such as code pages) is simply removed from memory and reloaded from disk as needed. Writable data is never removed from memory by the operating system. Instead, if the amount of free memory drops below a certain threshold, the system asks the running applications to free up memory voluntarily to make room for new data. Applications that fail to free up enough memory are terminated.
https://developer.apple.com/library/content/documentation/Pe...
Is there any way to tell the OOM killer which program to kill first?
From TFA:
>Without swap, the system will call the OOM when the memory is exhausted. You can prioritize which processes get killed first in configuring oom_adj_score.
The linked solution document is only available to registered RH users, though, and the name is actually oom_score_adj and not oom_adj_score.
`man 5 proc` has details, but tl;dr is set /proc/<pid>/oom_score_adj to -1000 to make a process OOM-killer-invincible.
The fun OOM analogy [1] that comes up when people propose different OOM killer designs:
> An aircraft company discovered that it was cheaper to fly its planes with less fuel on board. The planes would be lighter and use less fuel and money was saved. On rare occasions however the amount of fuel was insufficient, and the plane would crash. This problem was solved by the engineers of the company by the development of a special OOF (out-of-fuel) mechanism. In emergency cases a passenger was selected and thrown out of the plane. (When necessary, the procedure was repeated.) A large body of theory was developed and many publications were devoted to the problem of properly selecting the victim to be ejected. Should the victim be chosen at random? Or should one choose the heaviest person? Or the oldest? Should passengers pay in order not to be ejected, so that the victim would be the poorest on board? And if for example the heaviest person was chosen, should there be a special exception in case that was the pilot? Should first class passengers be exempted? Now that the OOF mechanism existed, it would be activated every now and then, and eject passengers even when there was no fuel shortage. The engineers are still studying precisely how this malfunction is caused.
By default, it'll start killing processes when free memory drops below 10%, though you can configure the threshold. I had the same problem for years, and then I started using earlyoom and I don't have to deal with it anymore.
A system backed by an SSD does degrade more nicely, though. The system visibly slows down but doesn't go to outright unresponsive like it does on a hard drive. You can make a case for letting that happen and having human intervention select the processes to kill, rather than letting the kernel do it. So, even though it still isn't really useful as an extension of RAM, it can still be useful in recovering from systems that you've run yourself out of memory on. Since putting an SSD in my systems I've actually gone back to running with some swap space. Though the fact I like hibernation sometimes is also a reason I run with swap in Linux on my laptop.
[1]: Swap will almost certainly completely blow out the buffers on those things and you'll be stuck with the raw hardware write speeds pretty quickly.
Use earlyoom instead of relying on oom-killer.
https://github.com/rfjakob/earlyoom
To quote from the description:
> The oom-killer generally has a bad reputation among Linux users. This may be part of the reason Linux invokes it only when it has absolutely no other choice. It will swap out the desktop environment, drop the whole page cache and empty every buffer before it will ultimately kill a process. At least that's what I think what it will do. I have yet to be patient enough to wait for it.
[...]
> This made people wonder if the oom-killer could be configured to step in earlier: superuser.com , unix.stackexchange.com.
> As it turns out, no, it can't. At least using the in-kernel oom killer.
And earlyoom exists to provide a better alternative to oom-killer in userspace that's much more aggressive about maintaining responsivity.
Both changes have made my computers much more usable. Systems should designed to fail fast when memory is low instead of slowing down.
That said, the article's recommendation was spot on in terms of making a conscious decision on how you want your system to behave when its coming close to running out of memory. Large swap space was originally the way you got those things that were too big to fit in memory to run, and now they are a way to essentially batch process very large data sets.
It would never actually hit the OoM killer, instead it would just lock up while it still technically had a few hundred mb of memory free.
From what I can tell, it was stuck in a loop evicting something from cache and then immediately pulling it back in from disk. Everything was technically still running, but the ui wasn't responsive enough for me to even kill a program.
Simply adding 200mb of swap would change the behaviour enough that the OoM killer would eventually run.
That was different in the early days, but that was because people accepted worse performance (GC that stops the world for seconds can be better than no GC, even when running a GUI).
Certainly nowadays, if you take out half the RAM, you will want to take out half the processes, too.
Now, desktops can have 32 GB of RAM, but everyone just uses it to run Chrome.
We write very memory-hungry software to make up for our copious amount of RAM.
.... which will happily chew away your 32 GB of RAM if you let it run for enough time :)
1) If you're ok with one machine dropping out of your system, you don't need swap.
2) You should never build a system where losing a single machine is a problem.
3) Therefore, you should never need swap
4) Perhaps there is an exception for a desktop machine, since it's doesn't fit rule 2.
A bit of a side ramble: Unfortunately, sometimes regarding rule 2, you already have a system where losing a single machine is a problem, and it will take time and resources to improve or replace it to the point where losing a single machine isn't a problem, so "in the meantime" you have to accept and support this.
Also, sometimes "the meantime" is very long. :-(
Also, by the time the system is improved to be more resilient, maybe you'll be working somewhere else or on something else, and, presto, you'll uncover some other horrible legacy system in your dependency chain that isn't resilient either. It seems as if at every organization that has had computers for long enough, there is an infinite supply of legacy systems.
Point being unless you only work with brand new things that themselves only work with brand new things, you can't get out of getting decent at managing services that aren't properly "any single machine can disappear" resilient
However, as pointed out elsewhere, if you're hitting swap your performance will be so bad you might as well have lost the machine.
A cluster of a few machines experiences a bunch of requests that trigger pathological memory usage. One machine OOMs, drops out. Now the rest of the cluster has to take more load, needs more memory, and increases the likelihood that the other machines also run out of memory.
A performance cliff (as you'd inevitably see while swapping) also puts you at risk of cascading failure. It might actually be better to completely drop out if the restart time is reasonably low. This is similar to GC thrashing with Java servers: many people prefer to configure their servers to suicide when GC time is over some threshold rather than try to go on as long as possible. I'm one of those people.
Better ways to avoid cascading failure are overprovisioning (RAM is pretty cheap for servers) and load shedding / graceful degradation at the application layer, coupled with care in client-side retry logic. (Avoiding accidental capacity caches, using exponential backoff on any retry.)
I suspect I would be fine with much of our datacenter being diskless (and put disk -- ahem, I mean storage -- where it is needed). Local disk is a headache more often than not.
The issue is not swap or swap utilisation, the problem is worst case latency. Even for a database an OOM is usually better than a latency hit that makes it unusable slow.
As a simple example an app might start allocating and use memory in an infinite loop. How long will that take? How long will your system be unresponsive?
If you have more swap than you can write in 30s you'll most likely do it wrong (your system can be unresponsive for 60+s).
Another worst case would be allocating all memory, using it and then performing random reads throughout the memory space. Your swap to ram ratio defines how much misses and thus how much IO you are doing instead of direct memory access. This should stay way below your IO capacity.
As a result I usually try to use a small swap partition and monitor for swap-ins, not swap usage or swap out.
So that's my thinking around swap mainly due to the fact that I have seen too many servers causing issues due to swap related latency.
In light of that this recommendation from Red Hat is very interesting. Just a fifth of memory as swap is probably enough to get real world performance back, without getting completely stuck when something goes haywire. On large memory systems it should probably be even less.
This sounds like good advice compared to the classic "2x RAM" guideline. Back in the HDD era when we already had around 8GB RAM I started wondering how long it would take to actually fill 16GB of swap in terms of raw write speed.
On the other hand SSDs are fast enough that swap might actually make a low-memory system feel faster.
My current Linux laptop has around the same amount of swap as RAM. Am I mistaken in thinking that suspend-to-disk saves RAM contents on the swap partition?
Even now, with not quite 9GB "in use" 22GB "standby" and 1G "free" the paging file is at 1.5% use with peak about 3%. Granted, that is tiny on a fixed 1.5GB file but for some reason Windows feels the need to drop about 20-50Mb in the pagefile.
I missed the mention of zram. Zram can create ramdisks, and compress them. It can create a compressed swapdisk in ram, basically making your ram last longer in case you really run out of memory. In my experience that is a good alternative to having a bit of swapspace as reserve, as the article recommends.
One neat thing that swap can be used for is to take sleeping processes, that might use a lot of memory, and put that memory on disk to free up memory for other active programs.
How much swap do you need ? Square root of your RAM size.
I shut it off because, in my opinion, OSX is pretty shitty at memory management. It swaps for no good reason. I've had 6 GB free and still using 1.6 GB of swap. That shouldn't happen.
Under Linux, if there is heavy swapping, forget it, nothing will work well.
In my experience, the most commonly used swap configuration is minimal, 500M perhaps. And vm.swappiness=1.
I'd say it's more rare to find a system that actually needs swap than one that can do without it.
I have yet to run into an application that for some reason needed swap around.
I got another taste of that lately when I had to build Wireshark from source on a Raspberry Pi model B thanks to broken packages in the repo. At least it was the version with 512MB of memory and overclocked. For the most part Wireshark isn't that bad to build, certainly better than Firefox, but there are a couple of dissectors that have unreasonably large source files.
it just slows my system down to a crawl, requiring me to force a reboot
it probably depends on your hardware
and if i disable the pagefile, windows update stops working and at 75% memory usage it starts panicing and closing programs
But it allows me to save, exit a program, or close a tab, without losing any work.
Question is: is there a way to identify who the culprit is and take appropriate vengeful action in the instance of swapping? Generally it's "foo large application", though there's also a very strong tendency for "foo large application" to be a critical system element -- either OS or application level.
1. Swap is slow
2. If using swap, your system starts to thrash
3. If thrashing, you can't close programs to free memory
4. If you can't close programs, you have to wait until the task is killed by the OS
5. If you have no swap (or very little), you don't have to wait.
Except with an SSD, swap isn't slow enough to cause that issue. So really this article only seems to apply to servers, not desktops.
Though it tends to mean you're boned, or going to be waiting a while while all i/o is dedicated to swapping for minutes at a time.
Of course the trend of making notebooks thinner by ditching the SODIMM slots and soldering insufficient amounts of memory to the mobo may reverse this.
Windows will.
It's a tradeoff: if you swap something out while not under pressure, that could be the thing you next need, resulting in it just getting swapped back in. Or, maybe not and the extra cache is useful (but if you're not under pressure, maybe letting go of older cache, instead of swapping something out, is a better trade: letting go of old cache doesn't require swapping out, since it's by-definition backed by disk.)
E.g., several large processes sleeping in memory on desktop would be fine if only one or two used at the same time. OTOH, clustered nodes well tuned for a single task may not need a swap.
In any case, it is a metric for thrashing that should be used to initiate culling.
[1] http://ubuntublog.org/tutorials/how-to-create-ramdisk-linux....
I have only ever enabled swap when compiling a few projects with "-j8" would take up all the memory. Using fewer threads would usually end up being slower than more threads + swap.
/shrug
Using Firefox on Bunsen Linux (Debian Jessie derived)
Its only using half of 4G of RAM, but there's 20 Mbytes in swap.
I've found myself wanting to upgrade it to 32G ram, but honestly that's about the only use case (besides production servers) where I would ever consider swap, and at that point I consider it a problem of not enough memory rather than swap being necessary.
Don't use swapfiles in FreeBSD because the file system write paths of UFS and ZFS potentially allocate memory. Both geom_mirror (software RAID-1) and geom_eli (disk encryption) are fine and I would recommend using GELI to create a onetime keys for mirrored swap partitions at boot.
An other good habit to get into is to limit the resources available to your services to some generous upper bound you expect them to require. The most flexible way to enforce those are restrictions in FreeBSD are hierarchical resource limits. Use them and monitor resource consumption. That way you get an early warning before a rouge process drives the system into massive swapping.