We're moving away from swap partitions on our Linux servers
utcc.utoronto.ca
utcc.utoronto.ca
If I created a swap partition in LVM I can force it and move it around to the best location WRT multiple hard drives and possible mix of HDD and SSD. With a file its a little more abstract and I could emulate it with some pointless abstraction, but why bother. A typical "this century" example would be running my boot and swap on local disk and storing all real world data on the NAS over iSCSI. Obviously now that "Everything" is virtualized this kind of system administration is void unless you're admin on the cluster itself LOL.
The problem with swap is memory seems to have gotten cheaper than storage and swap of any form takes storage space. Its not the capex that's the problem, its the opex of now you've got an extra 32GB that "has to" be backed up and "has to" be virus scanned or whatever security theatre, and that swap file "has to" be treated as the highest security PII HIPPA PCI category because who knows whats been swapped out of memory onto it, it could be chock full of CC numbers how do you know for sure unless its empty or not there? Paying for more memory doesn't increase your backup storage / security risk / virus scan times and the latter costs more than the former. I don't think non-cloudy admins realize there are cloudy admins out there with forests of little systems with like 8 GB of ram and 16 GB of storage, so the "old" rules from the 1980s about provision twice your ram as swap would increase the disk required from 16 GB to 32 GB doubling my costs for, essentially, nothing. Another problem with swap is in the old days if you needed more memory than you have physically you had to buy more memory, then for awhile you could emulate memory very slowly with disk but much faster than calling an IBM CE and buying more memory, now if you need more memory you spin up more instances or rebuild the instance on a larger flavor about as fast as how you'd add swap in the olden days, probably faster and it can be done dynamically in some architectures.
I've never thought about the AV scan of such a file!
As side note, not nervous when resizing NTFS in Windows, even for shrinking (still recommend to have backup for shrinking though).
Why on earth you'd backup swap?
Also encrypted swap is pretty easy to setup, then again most distros don't give that option on install.
We have tiny swap, like 1GB, working basically as "early warning" of "hey, this server probably could use few GBs more of RAM". There is very little use for any more swap than that
>I don't think non-cloudy admins realize there are cloudy admins out there with forests of little systems with like 8 GB of ram and 16 GB of storage, so the "old" rules from the 1980s about provision twice your ram as swap would increase the disk required from 16 GB to 32 GB doubling my costs for, essentially, nothing.
The "old rules" for swap were never relevant. But dumb myths die hard.
They did make sense in some contexts. Specifically for laptops expected to hibernate you need to preserve the whole ram. Then you need the extra size for whatever your swap was already using. For small ram systems, 2x the memory size was a good rule of thumb.
It stopped making sense when we stopped using swap and filing memory in normal situations.
I just yesterday had to re-partition a FreeBSD kernel test box with 256GB of RAM and a 2GB swap partition because I desperately needed a crashdump and the mini dump would have been 7GB.
The quotes around "has to", and especially the virus scan, makes me think there's a blanket corporate policy that applies to all files, regardless of file type.
Seems like a bad reason to disable swap, but I can't find a way to do what I really want, which is simply to reserve some CPU for input. And then I suppose in order to do something useful with it, require that any process that wants it be allowed at least some small allocation. Maybe that's the hard part and why I haven't found a way? It doesn't happen often, it's just really annoying when it does.
It should be possible to allocate say 95% of the system resources to the default cgroup and then you could create a secondary cgroup — recovery — which has access to the last 5% of the system resources which you could use to run commands such as “kill” or “top” to recover the system state.
Additionally you could run a second ssh server in the recovery cgroup which you could ssh into in the case of system lock up.
In reality it is probably easier to just reboot in most cases, or if you are dealing with servers use ipmi.
On Windows you don't have oom at all because it trades that solution for just swapping forever until either you can't do anything or manage to kill the right app yourself.
People have reported that their machines with small amount of RAM are now fully usable where previously the system become completely unresponsive when swapping started.
If it had some concept of 'in-use', for which you could define rules like 'has an active window' or 'is playing media', that might work better for me.
I have switched to systemd-oomd, but haven't yet gotten in any notable scrapes, so can't comment yet on how it fares.
Speaking of Windows: How is it that Linux seems to grind to a halt when it is forced to swap, while Windows works just fine in the same scenario? Sure, on Windows a lot of swap file "use" is just bookkeeping of empty pages because Windows doesn't overcommit memory. But even if it actually swaps it stays responsive, while I've had to reboot linux boxes more than once because they stayed unresponsive after memory pressure. Does Windows have a smarter swap algorithm? Is it more proactive about swapping stuff back in once memory pressure is gone? A better scheduler that gives more precedence to "interactive" things? And most importantly: can linux implement those things too?
/proc/sys/vm/swapiness or something like this. AFAIK shall be bigger than 50% to not see the behaviour that you describe.
Freudian slip?
Famous, huh? Do you have any resources or articles where I can read about this famous phenomenon?
Addressing your links in order:
1. I know that the word is thrashing, I was quoting the OP.
2. The existence of someone who thinks they experienced thrashing (but didn't do any investigation to see if they were actually experiencing thrashing), 13 years ago, doesn't seem very difinotive to me.
3. Please read this page. You'll see that the first sentence is "hard drive is being continuously thrashed even when no applications are apparently running". How is the system thrashing when there are no applications running? FYI, "thrashing" does not just mean ANY frequent hard drive reads for any reason.
4. I don't know what this article has to do with the question I asked. The article title is "How do I tell if my Windows server is swapping?" which seems completely unrelated to my question about how often Windows swaps?
Linux assumes it's better to have not-used thing in swap while RAM is used for IO caching, that's why you see this behaviour, there is (AFAIK) no mechanism to go "okay, we have few gigs free, let's bring stuff from swap back".
Reboot-less method is to add new LV with swap and swapoff the old one but that's more PITA than restart...
I've not done much in this realm since NVMe SSD has been common so not sure how this behaves these days.
So, what's the function of swap for Windows?
Overcommitting is a separate idea, which can be done without swap.
Therefore, we cannot say that "Windows doesn't overcommit". Windows applications can overcommit. You don't know which ones might do that, for what purpose.
It's the use cases that make a difference.
1. Makes claims about performance 2. Provides no benchmarks
The only concrete answer to "why have swap if I have enough RAM" is "well, if you run out of RAM..."
The only time swap isn’t just delaying (and growing) the inevitable shitshow is when, for some reason:
1) some program is consuming some significant amount of Ram, but doesn’t need to run for awhile, and it’s somehow faster/easier to let it swap to disk than shutting it down and restarting it while you do something else that consumes Ram.
2) something you’re running has some amount of bloat or leak that you know will never get exercised, but also never will grow larger than your swap file and take your system out entirely.
Both of those situations are not only rare, but easy to misjudge and end up wedging your system.
And then compilation eventually finishes and the allocated memory is released on process termination.
Disk cache is really flexible at the OS level. Look up 'vmtouch' on Linux and play with some recently used files, you'd be surprised how many are in memory. If you don't have swap and only have 'X' amount of RAM free the program will still load and run fine, it'll just do so in a way that more directly thrashes the disk. If you add some swap the OS will page out some other memory that's not being used giving 'X+Y' RAM free for disk caching that the OS will use.
In fact this usage of free memory as disk cache is so ingrained that memory used this way doesn't even appear allocated. OSes are pretty much all designed to use whatever free RAM is available as disk cache. The disk cache shrinks if that RAM is actually needed.
Effectively a great way to think of swap is that your OS is always swapping, even without a swapfile! If you read a file there's swapping of that file into memory, it's even done using the standard memory paging that swap files do. Without a swap file you still have swapping but you've just removed the ability for the OS to write out memory that it thinks is less relevant than the files you're currently trying to read. Which is almost always a loss.
Technically a program, but not what I think anyone typically considers one in this context?
What other ones are you aware of?
I ask because generally programs use RAM because it’s fast - otherwise they’d use something else. Having a program intentionally allocate more than is available would almost certainly result in serious performance issues it would have a hard time dealing with predictably - as compared to using disk for the ‘excess’, and then reading/loading what it needs.
In terms of what benefits from that most modern programs are actually surprisingly flexible in memory allocation. Anything with a garbage collection behind the scenes will flush more often if needed. Having swap to push the non needed programs out let's GC based programs keep allocations longer which is intended to enable reuse
Some just allocate when they need it until it doesn't work (cough Chrome cough), but that is a far, far throw from intentionally allocating more memory than the machine has, and none of them seem to do it to intentionally only work on a small subset at a time, even if de-facto that is what happens.
What sort of performance or reliability would you expect anyway from a database that intentionally loaded the entire database into memory first without even caring how much would fit?
Every database I’ve ever dealt with has fixed memory limits that get set.
Otherwise the database server will OOM even with swap if it doesn’t limit memory consumption and how much it loads, on any machine, with a given large enough database. And the size of the database is up to the user.
It’s a fundamental part of the problem.
Speed and reliability to some extent are literally core requirements of database servers, so any database software that doesn’t do it is going to have a bad time.
Linux with swap available can swap out never used pages (say a part of library that never gets called) and use freed memory to cache disk IO. We see that on every server that has enough data to fill the remaining RAM with page cache. But you only need like a gig of swap to take advantage of it.
The OS will page in files that are being accessed regularly. Without a swap file you don't stop the OS swapping between memory and disk, that'd be pathological since disk access is that much slower.
Instead what not having a swapfile really means is that you are telling the OS "anything allocated ever on the system is more important than disk cache".
It is working fine on Linux, you just have to set it up correctly.
> Reading from it returns the current image size limit, which is set to around 2/5 of the available RAM size by default.
I have successfully been hibernating a laptop after a few hours in suspend with swap allocated 50% the size of my RAM, so 8GB when I had a 16GB laptop, and 16GB now on a 32GB laptop. Never had any issues.
As a Java shop, the last thing you want to do is swap. Garbage Collection and Swap tend to be very bad cohabitants.
The only real thing that you may be missing is SSH being swapped out and some other system processes but those may be 2MiB altogether.
Maybe an option would be disabling swap just for the JVM. IDK if that is something you can do with cgroups. But if 99% of your anonymous RAM is pinned anyways there is little to gain.
The short answer is that not having swap causes the system to run with one hand tied behind its back, since it’s forced to store unused/rarely used stuff in RAM, taking up space that could be used more productively (like disk cache).
So the only use case I see for swap in zram is if you never want to swap to disk, but still want to use swap for some reason.
I’ll definitely double my RAM for my next system (in a year or two) and will definitely have to consider swap on zram then.
Incidentally, Darwin creates, defragments and reclaims swapfiles dynamically. All this complaining about swap partition / file sizes is kind of … silly. None of that code is rocket science.
One can always use and write wasteful programs but it is a choice, rather than necessity as folks often portray.
I used to do this as well, but consider that the VFS cache basically does the same thing. The main difference is that cached data can be dropped or flushed if necessary to free memory, whereas this can't be done automatically with tmpfs.
Benchmark both cases and I'll bet you'd be surprised at how little difference there is.
I do a lot of builds on my system and most are small enough to entirely live in RAM. But when you build Firefox it is best to start flushing some data to disk.
Entrusting this to the VFS cache has better predictability, and nearly the same performance in many cases.
Zswap is somewhat similar and might also work well for you, it uses disk swap but with a compressed cache in RAM.
At the beginning of the HDD was where the sectors are most rapidly accessed physically.
For SSD there should not be much difference between a separate partition or a file, or whether the reserved sectors the swap occupies are at the beginning of the drive. Should probably be aligned with the block size of the SSD though.
Still may be helpful to have a separate SSD for swap besides the drive the OS is on.
It would be worth trying but I like a swap file in Linux so I can confine Linux to a single type EXT4 partition right next to my NTFS WIndows partition.
With Windows it does work good without a swap file as long as you have plenty of memory left over after Windows and your desired apps are loaded. Handling further data loads within reason.
$5 per YEAR for a container, 128MB RAM, 1GB persistent storage, SolusVM user panel, and 123.45.67.89:16300 - 16320 NATd to your enp0s3. swapon/swapoff don’t work. I tried installing QEMU, didn’t work either because storage is I/O rate limited. IIRC, OpenVZ has some sort of unified cache control for RAM and storage I/O, and trying to circumvent quota through “disk” don’t work because of it.
I’m not sure what the use case might be, legal or not, but there seems to be dozen or so operators taking our credit cards in the exact same manner. Maybe the operation itself is a money laundering, or maybe it’s to run elaborate offsite contraptions for illegal activities, or to run SEO fake sites? I don’t know.
Even with openvz/LXC though, they will generally give you a few choices of distro flavors. So in general if given a memory allowance 256 or under, definitely don’t choose anything Red Hat based or you’ll suffer from yum potentially not working.
On cloud machines, LVM doesn't make a lot of sense, and neither does lots of partitions. Partitions are from a time where you have one disk and want to setup filesystem boundaries. But in a virtualised setup you can slice block storage any way you want, and present it to the VM as completely separate disks. It's a bit like having LVM "on the outside" of the VM. This way, the VM itself becomes simpler and has to care less about the specifics of the disk.
If you are on physical hardware, I'd argue that if you're not using LVM or ZFS, you're probably doing it wrong. Only really small scale (SD card, eMMC, USB SSD, raspberry PI type of stuff) works better without it. That said, even LVM wouldn't be too bad in such cases. (You'd have to get down to JFFS2 and the likes to really not use LVM)
If course if you need more than a bit of swap there is something suspicious about your system but I have often found this extra RAM is useful. For example I often have many open files in my editor, great candidates to swap out the ones that I am not currently editing. Or /tmp in RAM, I don't need it persisted but if the RAM could be better used for other things feel free to swap the unused files out.
I've been saying this for awhile, and I always get berated by angry linux nerds who didn't do any research. The times they are a changin'?
Thanks for the heads-up.
> (If you're willing to use swap files anyway you can always add more swap space with a swap file even if you started with a swap partition. However, if your swap partition is too big, shrinking it or limiting how much of it gets used is more annoying.)
You can't shrink swap file without swapoff & swapon anyway so it's illusion of an improvement
I didn't find it necessary to blog about it.
I guess NVMe gave it a second wind?
(The other argument is the usual mix of "but it can swap out unused areas" in which case you should fix the program to not allocate a ton of memory it doesn't need; and "but the OOM killer will come" and yes that's the entire point please move my process elsewhere).
Isn't this basically asking every application to independently reimplement their own swap-style functionality? Surely it makes more sense to have swap as a system level feature that every application can take advantage of?
That way the system has more information available to make swapping decisions since it can take into account the memory usage of all applications and drivers/etc together, and applications can take advantage of the most advanced swapping algorithms in the latest versions of the system without having to be individually updated.
No, it's asking for applications not to allocate a bunch of space they don't need in the first place.
In practice I expect the typical scenario you're thinking of is applications are allocating a bunch of memory which they use only briefly or rarely and could easily fetch or recompute that data again later if needed. That's exactly the scenario which swap optimizes for. So why should every application individually implement logic to optimize for it? Fixing these kind of issues is not as simple as "just don't allocate the memory".
That is what I'm suggesting, and I suggest you look at how much space is wasted by startup-initialized data in libraries for features you'll never use, JIT representations that never get compiled for more than one shape, class metadata that never gets touched and vtables which don't get used, how easy it is in any GC language to accidentally keep a large buffer alive until the end of a lexical scope rather than its last real use, etc. etc.
Today's programs are bloated as hell dude.
Especially if it’s a library or the like?
That infrequent call might be key for program stability in one context or pointless (garbage collector whatever- key in a major server program, pointless in a toy app).
The moral of the story is, if under load code or data gets moved from where access is fast (RAM) to where it is slow ( even SSD and most NVMe counts here) then gets accessed when under load? It makes the problem much worse.
Hence why the advice for servers is turn swap off. It’s a footgun waiting to happen.
Once you exhaust the RAM and you run with no swap configured, what happens with your application?
That rarely used code path will now take 50 seconds because it got called when it was swapped out (and the system is under memory pressure), instead of either OOM’ing awhile ago or completing in a couple microseconds like it usually does.
And ‘infrequently called’ here could be every couple minutes.
No, swap makes it worse. In virtually all cases I’d rather die than stall or flap.
I’m sure someone here will chime in with their example though.
I rarely need or set up much swap on a server, but I've managed to have the same kind of problem with the OOM killer not kicking in very quickly on a server. And on my desktops swap is able to soak up many gigabytes and improve my performance a lot with thrashing basically never happening.
I had it happen recently that the backup/sync software for a NAS had a constant factor memory consumption based on the number of files it was syncing. Transferred in a bunch of data, and blam. Commercial NAS, and they used swap (shitty, never buy QNAP). Wedged so hard it took a hard power cycle to even get console, AND caused data corruption in the ZFS pools.
It’s what happens next that decides the stability of the system.
If it can grow into swap and keep going, things start crawling, load builds up, buffers expand, and the system eventually grinds to a halt. On Linux, often in a really wedged and irritating/impossible to fix way.
Or, if no swap, OOM killer shoots something in the head (hopefully the offender), and we’re back to normal (minus the thing it killed, which is usually the thing you didn’t want it to - but at least the system works and you know something is wrong).
In either case, caches have dropped to near zero awhile ago so performance is already getting bad. It’s if we enter a death spiral, or get death early enough the whole system doesn’t spiral.
Right. You can get pretty deep into a death spiral even without swap. I feel like the better-performing solution is to make OOM trigger earlier, and to then go ahead and have swap be on.
Somewhat tangentially, I'd really like a setting for minimum disk cache, and that would do so much to help prevent thrashing.
(Which is why most big server deployments disable it.)
- swap on spinning rust outright kills the system if memory is low. it's a full blown crash.
- swap on ssd act's much the same
- zram swap stalls and slows down the system for several minutes
I just raise vm.min_free_kbytes to make the oom killer faster and disable swap. At least the machines don't crash hard this way.
Me and my team removed swap on tens of thousands of machines cause we were sick of dealing with this. We wanted the machines to fail hard and fast and not go into a state where it’s doing only a minuscule amount of actual work while trying to recover.
On most production systems, memory allocation is roughly 1% system and 99% application (not counting temporary, evictable allocations like disk cache). Modern applications are not designed to have their pages swapped out to disk and it is not particularly helpful to swap the teeny tiny bits of OS.
There are parts of this article which are true, but the overall picture it paints is not accurate and the conclusion that swap is helpful is not correct for high performance systems - both workstations and servers.
Doesn't have to be big, but without it, you can bring the system to its knees with 30% ostensibly free ram.
I don't find the term "performance" useful outside of a specific context. The main reason we don't use swap is predictability.
We know how much memory a particular VM should be using. If it exceeds that, failing quickly rather than changing behavior (slowing down) is far preferable, and then you correct whatever the problem is.
All the testing we've done with swap has shown it to have negative in server environments. I see it as useful for client machines, and maybe a bandaid to get by with underpowered systems if you have to for some reason. But that shouldn't happen in prod, especially if you're a public cloud user, as most folks here seem to be.
If your working set is bigger than your RAM you have a problem, swap or not.
However swap helps optimize your RAM usage so that your working set can fit with less waste. It allows the kernel more flexibility with what pages to evict from RAM. Without swap it can only evict pages backed by files. If you start evicting files that are in your working set you are just as screwed as if it starts evicting anonymous pages in your working set. With swap it can evict other unused pages before touching the pages in your working set.
I think there is some truth that it can be harder to notice with swap because there is more buffer between running great and literally crashing. However, in either situation you should be monitoring IO wait and application performance to ensure that you have enough memory for your working set.