Improved default settings for Linux machines
tobert.github.io
tobert.github.io
The mmap, file-max, and SHM advice is application dependent. Understand what your system is doing, and only increase if necessary. i.e. PostgreSQL < 9.3 is the only large user of SHM I can think of off hand.
The limits.conf advice is also bad. You should have a safety net here and increase these as needed and per user in /etc/security/limits.d
A less harmful guide would be something like "these are the knobs you may need to turn for certain apps, and here is the documentation on what they affect" - this looks a bit better https://wiki.archlinux.org/index.php/sysctl
Sysctl variables in Linux meet all these same criteria. Linux is already tuned for general use (or is damn close to it). Any knob you tweak is likely to make things worse in a way you don't understand. Just leave things unless you have a specific use case that requires different tuning.
The optimal value for what? Speed? Performance? Comfort? Economics? I don't drive a car, but wouldn't trade-offs apply to car tuning just like it does to everything else?
http://tweaked.io/guide/kernel/
In practice the file-handles tweak is the one that many people seem to forget, and can make a big difference for heavy webservers.
Side note: noatime is possible on OSX too [0]
[0]: https://github.com/tlvince/noatime-osx/blob/master/com.tlvin...
https://access.redhat.com/site/documentation/en-US/Red_Hat_E...
eg when I run it on one of our KVM hosts.
$ tuned-adm list
Available profiles:
- throughput-performance
- laptop-ac-powersave
- virtual-guest
- latency-performance
- enterprise-storage
- default
- spindown-disk
- desktop-powersave
- virtual-host
- laptop-battery-powersave
- server-powersave
Current active profile: virtual-hostI wonder if they made any new improvements in RHEL 7?
Easy enough to implement own scripts using the profiles as a basis.
Swap isn't some artifact from the days of 640k, used only because memory is expensive. Shit is always stored on disk; swap just allows that shit to be unused pages of active programs rather than actively used pages of files on disk.
Without swap, you force the kernel to prioritize cold code paths of rarely used daemons over, say, your web browser's cache. That's just dumb.
Every reference I've read suggests there's no need for swap space once you have more than ~2GB of RAM, but I find that extremely hard to believe.
When somebody asks this question, he always thinks this way:
- I have a system "A" with x GB of ram, with y GB of swap. - I have a system "B" with x+y GB of ram, and no swap, because it has all the virtual memory A is using.
well, the problem is that one should not compare system B with system A, he should compare system B with system C:
- I have a system "C", with x+y GB of ram, and z of swap.
system C will perform potentially better than B.
The generic explanation is that the kernel may decide that it's better to swap out some data, and use the space for caching purposes. This is a concept, that within limits, it's not related to the amount of RAM.
Edit: Possible caveat: Maybe whatever was running overnight did a lot of disk I/O and the OS decided to cache it at the expense of moving idle processes to swap (not sure if Linux does that or not).
In my experience (Ubuntu 9.10, kernal 2.6.31), with the default settings it does not do this during intensive disk I/O unless there is very little free memory to start with.
In other words, if the system is not low on memory to start with, I can start a bunch of processes that read and/or write multiple terabytes of data and when I come back the next morning little or no additional swap will have been consumed.
People should of course learn about these setting as they change them, but many of the suggested values really are improvements for lots of use cases.
I do wish people were more rigorous about changing swappiness. Maybe if swappiness=0 was documented as "pageout every byte of /usr/sbin/sshd and /lib/libc.so before you pageout a single byte of java's bloated heap" then people would be less eager to apply it.
Java's bloated heap, for example, has easy to use startup options that cause the Java process to use a deterministic amount of RAM, as does any other sane program that uses nontrivial amounts of memory.
OOMing (or swapping) a server almost always indicates an error that needs human intervention, so cleanly killing/rebooting the faulting system and raising the alarm in your monitoring system is the right thing to do.
Modern systems are usually teetering near the edge of using all physical RAM. This is by design; unused RAM is wasted. You want to use as much RAM as possible to avoid going to disk. What you call "OOMing" is when the applications on the system require more physical RAM than what exists. This is independent from the swappiness ratio.
The goal of virtual memory is to apportion physical memory to the things that need it most -- keep frequently used data in RAM, and pageout things which aren't frequently used. This makes things go faster. When you disable swap, you constrain VM's ability to do this -- now it must keep ALL anonymous memory in RAM, and it will pageout file-backed memory instead. Even if that file-backed memory is much hotter.
I came to this conclusion by following the kernel community, starting with the "why swap at all" flamewar on LKML. See this response from Nick Piggin http://marc.info/?t=108555368800003&r=1&w=2 who is a fairly prominent kernel developer. Nothing I've read from the horses mouth has refuted this since then. This is true even on systems with gobs of memory.
You're worried about systems which grind to a halt under memory pressure, which is unquestionably a concern. The thing is, disabling swap doesn't fix this. As soon as you're paging out important file-backed pages (like libc), your system is going to grind to a halt anyway. And disabling swap can't prevent this -- (same point from a VM developer here http://marc.info/?l=linux-kernel&m=108557438107853&w=2). To really give your prod systems a safety net, you need to (a) lock important memory (like, say, SSH and libc) in memory or (b) constrain processes which hog memory with (e.g.) memory cgroups. IMHO cgroups/containers are a better solution.
Ideal tuning would probably also reserve some decent amount of space for file caches and slab but I'm not aware of any setting that does that.
It's an entirely different story on laptops and development servers, where the workload varies widely, may contain large idle heap allocations worth swapping, and manually configuring memory usage isn't practical.
There is a school of thought which thinks especially server devs should just trust the operating system in this regard. One notable person of that school is phk of FreeBSD and Varnish fame: https://www.varnish-cache.org/trac/wiki/ArchitectNotes
But there are cases where it clearly does more harm then good. I had a postgresql datbase server with a lot of load on it. The server had loads of ram, more than what postgresql had been configured to use plus the actual database size on disk. Even so linux one day decided to swap out parts of the database's memory, i assume since that was very rarely used and it was decided that something else would be more useful to have in memory. When the time came for queries that used that part of the database, they had a huge latency compared to what was expected.
Maybe i'm misremembering and maybe there was some way of preventing that from happen while still having a swap enabled on the server..
Either way you handle it, the server is toast. So the real solution is to start dropping requests or downgrading service in some way to make sure the server never reaches OOM.
It definitely is important that the person signing the cheques knows how much capacity they're paying for, but that goes two ways. If you underprovision, they must understand that if they ever get mentioned in the NYT, their site will likely go down. It's up to them to balance the risks.
There's basically no reason not to have at least a small amount of swap. There's never been a benchmark showing no-swap as faster.
It should also be mentioned that things like mmap()+MAP_PRIVATE on a file will flat out fail if the file is larger than free_memory+swap. It's a common, easy and fast way to work with large files where the portion of the file you're working on gets paged in and out as required. Turn off swap and you break this functionality. You're basically limiting what you can do on your system with no performance gains.
Even with java apps (which I grant you are difficult to estimate memory limits on without knowing the application's design) swapping can be a useful way to page out unneeded parts of memory - and more importantly, keep needed parts of memory intact. This can also mean keeping the Java processes in memory while paging out apps which are less crucial, which leaves more room for Java, etc.
Swap is, for lack of a better comparison, the canvas sheet you use to catch someone jumping off a building, or a new york city sewer. In the first example you can use it to save your applications/servers so you don't need to reboot them (the higher your availability requirements, the less you can stand random reboots). In the second example it's the place you send those inhabitants you don't deem worthy of RSS.
The other thing you have to consider is the idea of memory overcommit in application and kernel design. The system is built with a promise that it has way more memory than it actually physically does. Applications will always reserve a fake huge chunk of memory and doesn't care that the system is lying to them about how much is really available. Without swap, when these apps attempt to use the nonexistent available memory, they crash. With swap, they survive.
Then there's apps designed to rely on swap like the Varnish malloc() and file storage methods, or database servers. Even if you disable swap there are still performance and stability problems related to the VM, and understanding this helps your apps run more efficiently. (http://blog.jcole.us/2012/04/16/a-brief-update-on-numa-and-m...)
http://techreport.com/review/26523/the-ssd-endurance-experim...
I've an SSD that's been running in my development Linux laptop for very, very close to three years. The drive houses a couple of encrypted swap partitions along with the rest of the system. According to the SMART attributes, I've written 18TB to it in that time.
Don't worry about SSD wear. Really, don't. Either you'll get a drive that succumbs to crib death or super-shitty v1.0 firmware, or you'll get a drive that will last until long after you outgrow it.
That said, on my Linode I have swappiness=1 and a few GB of swap on the SSD since it is memory-constrained.
In any case, I've been running vm.swappiness=0 or 1 on nearly every machine I touch for many years now and have yet to see any problems. Swap is almost always a bad idea on modern machines.
Thanks for reading!
I currently have 1.2GB of swap in use on a machine that is not doing any active swapping at all; that's 1.2GB more space for caching.
[edit for spelling]
If you ran out of memory and malloc returned NULL that would be better, however a lot of applications rely on overcommit so there is no good answer:
* you can run with RAM + as much swap to make all applications happy and overcommit off, and accept the long delays caused by swapping
* run with just RAM, overcommit on and no swap, and accept that your applications may be killed by the OOM killer anytime
This is independent of applications whose working set size exceed that of physical memory.
If you only run applications you wrote yourself, you can prevent this behavior.
Except the downside of turning swap off is that you (possibly) don't cache your disk as aggressively, but the downside for having too much swap is that ill-behaved programs can grind your entire system to a halt for minutes/hours at a time by having too big of an active dataset. We're talking multiple seconds of latency, where you can't get anything done.
For me that's way too big of a downside for gaining a few extra MB of disk cache. I'd rather have it OOM right away.
96 GB would be plenty of memory for what I need to do--mostly I benefit from lots of filesystem caching. BUT--there's another user of this system and he does most of his work using Matlab. For the stuff he's doing, Matlab routinely sucks down 10s of gigabytes of memory. And it may stay that way for days at a time.
Without swapping, Matlab will happily sit on all that memory indefinitely, whether it's doing anything or not. Meanwhile, everything I do takes ages because of the paltry amount of memory left for FS caching.
With swapping, I can get some of that memory back for FS cache when Matlab isn't being used, and it makes a huge difference.
There can never be enough RAM. ;)
I can't wait to see these settings cargo-culted onto systems of customers who then complain Linux doesn't behave the way they expect it to.
Next time, keep your sysctls to yourself.
How disappointing.
Some programs that use select(2) are known to assume that FD_SETSIZE is at least the maximum number of file descriptors available (instead of checking FD_SETSIZE). This lack of bounds checking may lead to a stack or heap overflow and a security vulnerability.
More recently, if you build with fortified glibc options, then you'll get automatic bounds checking, but do you know that your own daemons are built this way?
This is an example of why it's not a good idea to arbitrarily change a list of default settings system-wide without understanding the implications. The defaults have not been changed for a reason; otherwise distributions would already ship with these changes.
References: https://lists.ubuntu.com/archives/ubuntu-devel/2010-Septembe... http://www.outflux.net/blog/archives/2014/06/13/5-year-old-g...
It's much easier to overlook than you'd probably imagine. I have seen apps serving hundreds of thousands of API requests per day that had the default settings. It's one of those quick changes that can have a big impact.
Then why is this posted at all, this isn't improved default Linux settings, it's settings some guy likes for some customized environment.
This is now the default in DragonFlyBSD: http://freshbsd.org/commit/dfbsd/3a877e444fff816b8a340d35fe3...
Also, I think I would prefer process sbrk failure to OOM killer activation. So setting vm/overcommit_memory=2, overcommit ratio to 80%, a decent swap size, and code actually handling errors. IE consistency versus randomness.
Not that randomness is bad for testing, cf Chaos Monkey: https://github.com/Netflix/SimianArmy/wiki/Chaos-Monkey
My philosophy is keep it do default unless you have an issue. Guess what ? It works just fine.
For example, just this morning Chromium started failing because I hand't disabled limits on one of my machines. I pulled down my standard settings, applied them, then the problem went away. It wont' be coming back either.
From unswappable kernel memory, yeah.
> kernel.pid_max = 999999
> * - nproc unlimited
1.000.000 pids is 1 mil task_struct's.
On my quite stripped out kernel 14 task_struct's fit into order 3 slab -- 14 objects per 32KB or kernel memory.
1000000 / 14 * 32 * 1024 = 2.18 GB of kernel memory
and that's not even counting other kernel structures!
As alluded to before, defaults are default for a reason. Having someone explain why they change them is a good exercise for both reader and author.
for example fiddling with swappyness means that you'll end up with less RAM for important things, like file cache.
If I use the --disable-gpu flag, I will hit the file handle limits. I have increased the limits and it works fine now.
I have an AMD GPU and chrome/chromium just does not work. It will constantly flicker.
sudo sysctl -p /etc/sysctl.conf
My oldish kernel doesn't recognize the PID settings, which is unfortunate.
Teh sound is so much more sound-ier! Way much more cranked up than the lame defaults! http://goo.gl/TJLTMF