The Linux kernel's inability to gracefully handle low memory pressure
lkml.org
lkml.org
`pthread_create` can sometimes return back a garbage thread value or crash your program entirely without any way to catch it or detect it [1]. High speed threadding is hard enough as it is, without the kernel acting non-deterministically.
Un-killable processes after copy failure (D or S state) [2]. If the kernel is completely unable to recover from this failure, is it really best to make the process hang forever, where your only available option is to restart the machine? I ran into this with a copy onto a network drive with a spotty connection, that actual file itself really didn't matter - but there was no way to tell the kernel this.
Out Of Memory (OOM) "randomly" kills off processes without warning [3]. There doesn't appear to be a way to mark something as low-priority or high-priority and if you have a few things running, it's just "random" what you end up losing. From a software writing stand-point this is frustrating to say the least and makes recovery very difficult - who restarts who and how do you tell why the other process is down?
[1] https://linux.die.net/man/3/pthread_create
[2] https://superuser.com/questions/539920/cant-kill-a-sleeping-...
[3] https://serverfault.com/questions/84766/how-to-know-the-caus...
Overall I ended up in the camp "you shouldn't throw passengers out of the plane"[1]. The best way to have the OOM killer behave well is not to have it run at all. If I don't have enough RAM just panic and let me figure out what needs to be done.
The current approach does indeed work much better. You can entirely disable the OOM Killer for a given workload with those procfs handles.
The way that the score is calculated, as I understand it, the processes with large memory footprints that are not particularly long-running, and especially not owned by root, are the most likely to be killed. My shared servers are running unicorn for Rails apps, so they almost all look the same under that lens.
I want the process with the runaway memory usage to be killed, so that person becomes aware of the problem when they find their app is down. (There are many better ways to solve this, but in our current system design, there's nothing to tell us that a service is down, so killing some other disused service means it's down until some user comes along and is impacted.) It seems like killing the process that made the last allocation is the most likely way to get this behavior I'm looking for, where the failures are noticed right away.
But I'm not convinced I've fixed anything, even if the behavior characteristics seem better, and I haven't had any servers going rogue and killing themselves lately, I am not the sysadmin, I'm pretty sure the sysadmin just comes and makes the machine over-provisioned so this doesn't happen again, which seems to be the best advice out there. And as I've learned from the discussion around this article, the fact of having a fast SSD on which paging stuff out and expiring it from the swap, all happens fast enough to look almost as memory, almost completely neuters the OOM anyway.
Our whole situation is wrong, don't even get me started; I'd like to solve this with containers, where we can set a policy that says "all containers must have memory request and limit" and thanks to cgroups, this problem never comes back.
cgroups doesn't save you from the OOM killer. In fact, we're seeing persistent OOM issues in production even when there's more than enough physical memory to satisfy all allocation requests.
Similar to the issue discussed in the leading LKML thread, I/O contention and stalls when evicting pageable memory (e.g. buffer cache) can result in the OOM killer shooting down processes when an allocation request (particularly an in-kernel request, such as for socket buffers) can't be satisfied quickly enough--quickly according to some complex and opaque heuristics.
The fundamental problem is that overcommit has infected everything in the kernel. The assumption of overcommit is baked into Linux in the form of various heuristics and hacks. You can't really get away from it :(
Yeah I know, hence the use of "random".
> The way to mark priority in killing is to adjust this score through /proc.
Haven't heard about this, thanks for the heads up!
While going down that road is technically correct, it is a road full of pain.
A slightly less painful strategy is to disable overcommit. That way, if memory pressure is high, and a process calls `malloc`, that call will fail if there is not enough memory, and that process will fail. If you only have a couple of processes in your system that are using most of the memory and you can control them, it is simpler to just making them resilient to these kind of errors, than to try to mess with the process score to control the OOM killer.
There are totally legitimate use cases for processes that use significantly more virtual memory than physical memory (since virtual memory is relatively cheap, but physical memory isn't). A lot of programs are going to touch and not release all virtual memory they allocate from the kernel, but there are plenty of important counterexamples.
fork()/exec() is one example (which I've been burned by personally), but there are plenty of others. Any program that uses TCMalloc and has fluctuating memory consumption will have a lot of virtual address space allocated but not backed by physical pages. Sophisticated programs like in-memory caches or databases can also safely exploit a larger virtual memory space while keeping the amount of physical memory bounded.
https://backdrift.org/oom-killer-how-to-create-oom-exclusion...
Part of the issue with processes stuck in D state (waiting for the kernel to do something) is that it is deeply tied into kernel assumptions about things like NFS, NFS is stateless, and theoretically severs can appear and disappear at will, and operations will keep working when it comes back. You can make NFS a hell of a lot less annoying in this regard by mounting it with soft or intr flags, however if the network disappears or hiccups, you WILL lose data (the network is NEVER reliable, in fact the entire model of NFS is arguably wrong to begin with)
That new request you just received when a hardware failure occurred? Say goodbye to it, you've no way of ensuring it will make it to storage when the disks have caught fire. Later on, when you've put out the blaze, all that algorithms can do is tell you when things started to get lost.
However, it scales like crazy and is very powerful and reliable.
On local networks (everything attached to a single switch) with good hardware, it is reliable, and soft/intr is the worse choice among others.
To wit NFS is one of the commonly supported VM storage options (libvirt, VMware, etc).
There are many things "known" about NFS that are legacy.
That said, it's still entirely possible I've made a mistake. Please see here: https://github.com/electric-sheep-uc/black-sheep/blob/master...
The idea is that you do preliminary processing of a camera frame before sending a neural network over the top.
Could be a memory fragmentation issue
EDIT: Also since you're using C++ threads: You should really use move semantics, because right now you have two points of failure acting on the same thing: `new` operator may fail on creating the threads instance, and the underlying pthread_create may fail as well.
In all the C++ standard library implementations of std::thread the only member variable of that class is the native handle to the system thread itself; there are no additional member variables! This means that the size of a std::thread object is equal to the size of a native handle, usually the size of a pointer, but sometimes smaller.
If you create std::thread by `new` you're essentially creating a pointer to a "possibly a pointer", which comes with all the inefficiencies associated with it: Double indirection, small size allocations tend to fragment memory. And at the end of the day to actually use it, you have to at least lob around that outside pointer around on the stack anyway.
So there is zero benefit at all of using dynamic allocation for std::thread. Don't do it! Just create the std::thread instances on the stack, they're just handled/pointers with "smarts" around them, and you can copy them around just efficiently as you can copy a pointer or an integer. Better yet, if you're not trying to "outsmart" the compiler you'll often get copy elision where applicable.
Creating and destroying threads in a tight loop (and blocking on their destruction, thereby reducing the point of having many) seems like a bad idea. Conceptually you have only maximum of two threads in any given part of this snippet, the one running the loop and the current instance of visThread. My guess is also that the loop thread spends most of its time waiting for the recently created threads to die. Why not have visThread only get created once and process a queue of events?
Anyway without additional evidence you potentially flagged a bug in the c++ standard library rather than pthread_create.
> Creating and destroying threads in a tight loop (and blocking on their destruction, thereby reducing the point of having many) seems like a bad idea.
The purpose is that it's pretty much 100% always running a visThread (it's a neural network that takes about 100ms per image). The pre-process on the other hand runs in about 10ms, but there's no reason why it can't be run in advance (+). The neural network can't really be run on multiple cores safely, but it does have OpenMP (parallel loops).
(+) It does create some latency in the output compared to the real world, but it's not a massive deal when it comes down to it.
> Why not have visThread only get created once and process a queue of events?
This is probably the best way to do this, but this was the lazy way with not too much overhead (I think) :)
> Anyway without additional evidence you potentially flagged a bug in the c++ standard library rather than pthread_create.
Possibly, I did run it through GDB and Valgrind, it reliably seemed to die in pthread_create, but that of course could have only been the trigger. It could also be the aggressive optimization [1].
[1] https://github.com/electric-sheep-uc/black-sheep/blob/0735de...
I would also love to know if anybody has a solution for getting video to play properly in Firefox. I know it's not a bug per se, but it would be nice to not have to switch between browsers all the time.
I've been using Ubuntu for about a year now and otherwise its been a very positive experience.
Unfortunately, I don't really have a solution for this. I have an SSD as the main disk and _even now_, when I hit this too hard Linux grinds to a halt. No mouse, no keyboard, just heat, fans and disk light.
One thing that sometimes works for me is the old CTRL+ALT+FX mashing, but not always. Once you can get a shell you can type into you're okay, but of course this doesn't always work.
> I would also love to know if anybody has a solution for getting video to play properly in Firefox. I know it's not a bug per se, but it would be nice to not have to switch between browsers all the time.
What do you mean? It's been reliable for quite a long while? There were two issues I used to have, one was not having my graphics card setup (it was running from the CPU) and the other was not having flash (when that was something).
> I've been using Ubuntu for about a year now and otherwise its been a very positive experience.
Yeah, I think it makes one of the better daily drivers.
Re: playing video, for some reason I can't play South Park on either Firefox or Chrome. Videos on twitter also won't load with Firefox, though they work fine with Chrome.
Yeah, it's another bug with Desktop based Linux. The problem is that Linux "basically" treats the GUI like any other program, when things get heavy everything gets roughly evenly screwed. OSes designed to be centralized around a GUI on the other hand usually guarantee that GUI related processes get a minimum amount of time on the CPU to ensure they don't freeze.
Linux should 100% be doing this. Doing lots of hardware I/O shouldn't mean you lose the mouse or keyboard. When you lose control of input, you think the machine isn't doing something, when in actual fact it's doing tonnes, it's just not showing you. Even something like Android suffers from this under heavy load, it's really crap.
The real joke is, it's probably a difficult kernel fix. You would need some kind of watch dog timer for the kernel to make sure it's not getting too bogged down with any one particular task and then interrupt ones that are (some tasks don't like to be interrupted) [1]. You then need make sure that all of the heavy kernel calls don't make guarantees about the call being completed (i.e. blocking), which to be completely honest should be the default position to take anyway.
I'm not completely up-to-date on this, but my bet is that the issues come from everywhere. Programs can read/write arbitrarily large amounts of data (RAM, disk, network, bus, etc), when I believe you should be able to ask the kernel what size it would like you to read/write based on I/O activity and the capabilities of the device. If your program is bogging down the kernel, it should lover your recommended block read/write size. Better yet, this would be compatible with existing software, as they could choose to ignore this, possibly with it "punishing" programs that eat up lots of kernel time by making them wait longer for their next opportunity. There's a bunch of algorithms for time splicing tasks, but the most optimal appears to be max-min with very little organizing overhead [2].
> Re: playing video, for some reason I can't play South Park on either Firefox or Chrome. Videos on twitter also won't load with Firefox, though they work fine with Chrome.
Hmm, that shouldn't happen, sounds like you potentially have some system-wide badness. A few things to try:
* Disable any customization you made (extensions, add-ons, etc) - see if it is one of these interfering
* Make sure you have all updates and you're running an up-to-date version of Ubuntu (this issues have possibly been patched already)
* Make sure you have the correct drivers installed for your GPU
But... One thing I did note was that JavaScript coin miners have gotten so bad that I can't run certain sites without ad-blockers anymore (uBlock Origin is generally recommended). I remember my CPU sitting at one core maxed out just because of the JS engine. I generally run uBlock Origin + NoScript on every page and manually enable temporary scripts on pages I trust. One of the biggest offenders for crazily heavy JS was actually Facebook.
[1] https://en.wikipedia.org/wiki/Watchdog_timer
[2] https://uhra.herts.ac.uk/bitstream/handle/2299/19523/Accepte...
>> JavaScript coin miners have gotten so bad that I can't run certain sites without ad-blockers anymore
Whoa, is that the reason why some websites are eating a ton of memory?? Is this common?
>> I generally run uBlock Origin + NoScript on every page and manually enable temporary scripts on pages I trust.
Thank you for the recommendations, I'll check those out.
This does depend on cgroupsv2, but it works on most modern distros.
[1]: https://www.freedesktop.org/software/systemd/man/systemd.res...
Few programs can handle a fail return from "malloc", and Linux perhaps tries too hard to avoid forcing one. Most programs just aren't very good at getting a "no" to "give me more memory" Browsers should be better at this, since they started using vast amounts of memory for each tab.
I used to hit a worse bug on servers. If you did lots of MySQL activity, so that many blocks of open files were in memory, and then started creating processes, you'd often hit a situation where the Linux kernel needed a page of memory but couldn't evict a file block due to some lock being set. Crash. That was years ago; I hope it's been fixed by now.
Browsers are quite good at this actually. Major web browsers run on Windows (and even 32-bit windows!), where there is no overcommit, so malloc can return "no" any time, which happens quite often when you are limited to 4Gb of memory per process.
The only apps that suck at this are Linux-only apps that are never used anywhere else and just assume that all Linux systems have overcommit enabled.
I suspect that overcommiting is one of the reasons for this though. Many programmers in the Linux world have integrated that "malloc can't fail" and the only error handling they bother doing is calling abort() if malloc fails.
Of course the fact that C doesn't provide any sane way to implement error handling probably doesn't help.
While C has no special error handling mechanism in place, error handling can still be done reasonably. IMO, the big reason for why malloc() errors are rarely handled is because it is quite hard to come with a viable fallback strategy.
True, and that makes sense for something like git. But in my experience many long-lived programs don't bother to handle ENOMEM gracefully either.
But I guess I'm veering off-topic here, I'm mostly fine with applications crashing of their own volition when they don't have enough memory. I agree with you that in many cases there's no clear recovery path for an application that's out of RAM. It's the OOM-killer I have a problem with.
>While C has no special error handling mechanism in place, error handling can still be done reasonably.
I very much disagree with that. There are a few factors that make error handling in C a pain:
- No RAII, so you have to explicitly handle cleanup at every point you may have to early-return an error (goto fail etc...).
- No convenient way to return multiple values from a function. That means that in general functions signal errors returning some special value like 0 or -1 (even that is very much nonstandard, often even within the same library).
Oh you want to be able to signal several error conditions? Uh, maybe use several negative codes then? Oh you need those to return actual results? Well maybe set errno then? Don't forget to read `man errno` though, because it's easy to get it wrong. Oh you had a printf in DEBUG builds in there that overwrote errno before you could test it? Oops. Don't do that!
What's that, your function returns a pointer and not an integer? Ah, mmh, well maybe return NULL in case of error? You want to return several error codes? Well maybe you can just cast the integer code into a pointer and return that, then use macros do figure out which is which. It's terribly ugly? Well the kernel does it so... It can't be that bad right? Oh and what about errno? Remember that?
What's that, NULL is a valid return value for your function? Uh, that's annoying. Maybe use an output parameter then? Oh, or maybe some token value like 0xffffffff, that probably won't ever happen in practice right? After all that's what mmap does.
So no I wouldn't consider C error handling reasonable in any way shape or form. "Non-existent" is more accurate. You can always work around it but it always gets in the way.
I try to always implement comprehensive error checking in my programs. I do a significant amount of kernel/bare metal work, so it's really important. It's not rare that I end up with functions that contain more error-handling-related code than actual functional code.
C error handling ergonomics are non-existent which means that everybody bakes ad-hoc library-specific conventions that are extremely error-prone.
You could argue that they're doing it wrong and you might have a point but if almost everybody gets it wrong maybe it's fair to blame the language itself a little bit.
bool parse(inp_type *a, out_type **b, out_error **c);
where the return value false indicates an error. In C++, you'd just have written something like: out_type parse(const inp_type& a);
and thrown an exception on error.I think RAII can be useful, but I've never found any use for it in systems level code that I write. Most of the time I'm dealing with resources that were allocated inside a systems library or an external component which just gives me a handle to the resource. I think this is a common enough scenario in systems code that I don't think its just me.
e.g.
1. X = CreateResource()
2. Y = TransformResource(X)
3. ProcessNewResource(Y)
4. Z = TransformResource(Y)
5. etc. etc.
And so as you transform that resource, you will have multiple ways to unwind the resource depending on where the failure occurs. Even if you wrap X in some RAII container, you don't know what your destructor is going to look like.Another con to RAII, especially when paired with shared-ownership smart pointers, is you lose predictability over your resource deallocs. You never know when the last pointer is going go out of scope, and if its a 'heavy' resource with a complicated unwind, you're going to get a CPU spike at an indeterminate time. I deal primarily with industrial automation code and I much prefer to have a smooth/even CPU graph. I think this issue is more relevant to systems code which is the context of this thread.
What underlies this? I am astounded to see 1GB of memory returned when I close a couple of tabs.
Chrome and Firefox both seem like this.
Or are there any browsers where you observe significantly less memory usage on the same websites? (Ignoring limited browsers like Lynx of course)
I believe to have read somewhere that at least Chrome listens to low-memory situations broadcast by the OS and will evict such caches. So while it uses a lot of memory as long as memory is available, it will also release much of it if necessary.
I wish there was some sort of allocation API specifically for allocating caches so that recently accessed files could kick out a web browser's cache of a not-so-recently accessed tab or vice versa.
Assumining RGBA it's just over 8MB for a 1920x1080 image.
If I had to choose one program that is proportionally used the least alone, I would have voted browsers.
Unpopular opinion time:
My guess is most developers don't care and are not even looking at this anymore, either during development or after release. Nobody seems to even know how much memory their program allocates and how quickly it allocates that memory under various running conditions. I used to challenge my fellow developers: Stop in the debugger right now. About how much memory should the process be using? Nobody even seems to have an order-of-magnitude guess anymore. It's your program, dude! Shouldn't you know this?
You can ask any embedded software engineer exactly how much memory his/her program uses, what's the stack size, what's the heap size, what's statically and dynamically allocated. Sadly, this discipline is pretty much gone outside of that specialized area.
That said, Firefox some strides starting in 55 insofar as handling very large number of tabs starting in version 55.
https://www.techradar.com/news/firefoxs-blazing-speed-with-h...
Unfortunately Palemoon is more or less firefox 38 still isn't it?
But my point was more that sans JS it really seems to use up far less memory. Honestly, try nuking js for 1/2 hour and see how it feels.
I have this in Matlab on Linux. Matlab can actually deal with worker processes being killed, but my machine just locks up. Therefore, we have to run these specific simulations under Windows, where this doesn't occur.
In my case it happens like this:
I have a long running PHP process that constantly fires away mostly SELECT but also a bunch of INSERT and UPDATE statements and also some DELETEs.
Since the DB and the key files do not fit into memory, its all disk bound work.
All tables are MyISAM.
Like clockwork, this stalls the virtual machine once per day.
All I can do is to hard power down the VM and restart it. Afterwards the table data is corrupted beyond repair.
Not sure it is related to memory though. Because the memory usage of PHP and MySQL seem to be constant. Most RAM seems to be used by Linux for caches.
A good way to catch this would be to have something log the list of running queries every couple of seconds. Look at this log after the crash and you'll hopefully be able to identify which are the long running processes, and which are the regular queries that builds up.
To fix it would be a combination of making the queries that cause the locking to be less like that, also perhaps putting in a limit on how many queries can build up and also implement a way for the regular queries that build up to time out or fail quicker or more gracefully.
That should not happen with a DB even if you turn off the power. Are you sure the hardware is good?
When the kernel is running out of memory it will just start the OOM killer which will kill a process with low "nice" value.
But... it's still no good saying 'make your program behave nicely when malloc fails' - even if your own code is perfect, what are the chances that every library you use does the same thing? And even then, Linux by default will optimistically over-allocate memory (and rightfully so!) - with the result that you'll never catch every out of memory condition.
IMO, 'out of memory' is not a property that each single process should try to manage, rather it should be the OS or some other process with a global oversight that monitors memory usage and takes measures when memory gets tight.
The other point is that there's a distinction between kernele's view of OOM condition and some memory managers's OOM condition. Consider you run two processes, both allocate X gigs of memory and both succeed. However once you start running and committing the memory you'll get a kernel OOM condition and one process is killed. This is the overcomitting you mentioned.
Personally I don't see why people make such a big fuzz about dealing with memory allocation failures. Memory is just a resource same as any other OS resource, socket, mutex, pipe whatever. Normally in a well designed application you throw on these conditions and unwind your stack and ultimately report the error to the user or print it to the log and perhaps try again later. Just because it's "memory" should not make it special IMHO.
That's very much not true when 32-bit processes are involved. You can easily be out of (non-fragmented) address space in a 32-bit process (whether it's all resident or not) while the overall system is nowhere close to being out of memory.
Even in a 64-bit process you can exhaust the address space without being out of memory if you try hard enough; you just have to try much harder.
That said, even on Linux allocators will return NULL when they're just out of address space; there's no overcommit going on there.
Fewer than should, that's for sure, but hardly a trivial number. A lot of old-school C programs are very careful about this, and would handle such a failure passably well. Unfortunately, just about every other language tends to achieve greater "expressiveness" by making it harder to check for allocation failure. How many constructors were invoked by this line of code? By this simple manipulation of a list, map, or other collection type? How many hidden memory allocations did those involve? I'm not saying such expressiveness is a bad thing, but it does make memory-correctness more difficult and so most programmers won't even try.
As the world moves more and more toward "higher level" languages, returning an error from malloc becomes a less and less viable strategy. Might as well just terminate immediately, since "most frequent requester is most likely to die" is better than 99% of the OOM-killer configurations I've ever seen.
I don't see why Linux couldn't do the same; open /sys/kernel/something and epoll on it.
During this - unlike Linux - you can actually use the mouse, CLI and close programs yourself.
On top of that server applications like IIS has built-in watchdogs. If an IIS process grows to use too much memory (60% IIRC) or excessive CPU, the watchdog will recycle the process.
1. Show a GUI with a choice
2. Show a message on the current terminal and ask what to do
3. Just return "kill it now" if there is no interactive session
And if there is no such service, just default to 3. The problem really is that the state cannot be captured and communicated to the user. I doubt the NT kernel itself shows a GUI window, it's probably a service that gets woken up by a kernel exception, which in turn shows the window. Basically, the Linux kernel needs more pluggable functionality for user interactions. It's absolutely fine and even recommended to not have an entire GUI in the kernel, it needs to just provide a mechanism for userspace to capture the event and decide what to do with it.
I just kinda assumed that's how computers worked until I got a Mac a couple of months ago...
The link suggests that there might be some default parameters you could change to protect against this behavior. Does anyone have any suggestions on what settings to change?
This is because the offending program has allocated a lot of private dirty pages, which can't be dropped from memory because without swap space, there is nowhere for it to go.
IMO despite the standout behavior, I prefer my software to deal with itself.
Systems designed to wait for user input end up having design choices intent on keeping a user using them.
Software is just a tool. Not a lifestyle. Set and forget this shit as much as possible
From that reply it seems that Facebook implemented something similar, I guess for their servers.
Even if it's a small percentage of the overall "computing" population, there are still millions of people running Linux on the desktop (roughly 2% out of 3.2 billion people using internet makes for 64 million - a large european country). It's 64 million of people for which this behaviour is a pain in the arse.
I think you are confusing the issue raised here with your desktop experience.
Many long-lived processes are completely idle (when was the last time that `getty ttyv6` woke up?) or at a minimum have pages of memory which are never used (e.g. the bottom page of main's stack). Evicting these "theoretically accessible but in practice never accessed" pages for memory frees up more memory for the things which matter.
This comes into play when you copy or access huge files that are going to be read exactly once, they will start pushing out untouched program pages to disk, in exchange for disk cache that is completely 100% useless, even to the tune of hundreds of gigabytes of it.
Programs can reduce the problem with madvise(MADV_DONT_NEED), but that only applies to files you are mmap()ing, and every single program under the sun needs to be patched to issue these calls.
You can adjust vm.swapiness systctl to make X larger, but no matter what, programs will start to get pushed out to disk eventually, and cause unresponsiveness when activated. You can reduce vm.swapiness to 1, but if you do, the system only starts swapping in an absolute critical low ram situation and you encounter anywhere from 5 minutes, to 1+ hour of total, complete unresponsiveness in a low ram situation.
There _NEEDS_ to be a setting where program pages don't get pushed out for disk cache, peroid, unless approaching a low ram situation, but BEFORE it causes long periods of total crushing unresponsiveness.
Here's the thing: a mapped program page is just another page in the page cache. Now, you could maybe say that "any page cache page that is mapped into at least one process will be pinned", but the problem there is that means that any unprivileged process can then pin an unlimited amount of memory, which is an obvious non-starter.
A workable alternative might be to add an extended file attribute like 'privileged.pinned_mapping', which if set indicates that any pages of the file that have active shared mappings are pinned. That means the superuser can go along and mark all the normal executables in this way, and the worst-case memory consumption a user can cause is limited by the total size of all the executables marked in this way that the user has access to.
https://www.suse.com/documentation/sles-for-sap-12/book_s4s/...
It is quite effective, although historically there have been issues with bugs causing server lockups in the kernel code around this tunable. It seems to be quite stable in SLES 15, however.
While the tunable is available in their regular SLES product, it is only supported in the "SLES for SAP". The two share the same kernel, that is probably why.
Nobody is suggesting these pages be pinned which is an extreme measure.
That might be OK on a single-user system but it doesn't fly on a multi-user one. That's why I suggested you could gate that kind of thing behind some kind of superuser control.
The kernel could still “fairly” evict pages across users - just letting them choose which N pages they prefer to go first.
fadvise provides the same for file descriptors. some tools such as rsync make use of it to prevent clobbering the page cache when streaming files.
heck, if I were still a phd student, I'd want to run performance numbers on this in many different scenarios and see how performance behaves. feel like there could be a paper here.
[1]: https://wiki.freebsd.org/ZFSTuningGuide#L2ARC_discussion
Whenever I take a backup of my computer it winds up swapping everything else out to disk. Normally I'm perfectly happy letting unused pages get evicted in favor of more cache, but for this specific program this behavior is very much not ideal. I'm asking here since I've done some searching in the past and not found anything, but I'm not sure if I was using the right keywords.
THIS. I ended up disabling swap because my kernel insisted on essentially reserving 50% of RAM for disk buffers; meaning even with 16GiB of RAM, I'd have processes getting swapped out and taking forever to run, because everything was stuck in 8GiB of RAM, and Firefox was taking 6GiB of that. I couldn't for the life of me figure out a way to get Linux to make that something more reasonable, like 20%. (And yes, I tried playing with `vm.swapiness`.)
> This comes into play when you copy or access huge files that are going to be read exactly once
Did you try different vm.vfs_cache_pressure values?
TBH, there are very few OSes that get high-pressure resource scheduling and prioritization right under nearly all normal circumstances.
The hackaround for decades on Linux is always adding a tiny swap device, say 64-256 MiB on a fast device in order to 0) detect average high memory pressure with monitoring tools 1) prevent oddities under load without swap (as in OP example).
I would have thought some of the IRIX scheduler made it into Linux by now.
https://en.wikipedia.org/wiki/Zram
https://wiki.gentoo.org/wiki/Zram
Or Armbian, a Debian derivative for many RaspberryPi-like ARM boards, where it helps avoiding trashing the usual sd-cards, and holds the logs, with different compression algorithms for /var/log, /tmp and swap.
Their script is here:
https://github.com/armbian/build/blob/master/packages/bsp/co...
I really like how this enables a tiny NanoPi Neo2 with one GB Ram booting from a 64GB SD-Card in an aluminum case with SATA-adapter and a 2,5" HDD mostly idling to draw only about 1W from the wall socket, while clocking up from about 400Mhz to 1,3Ghz if it needs to. It's not the fastest, but sufficient.
But if it solves problems, why doesn't the kernel automatically assign part of RAM as a virtual swap device, if the user allows it. It can help monitoring tools, and because the kernel knows it is not a real swap device, it can optimize its use.
Linux does unfortunately have serious issues to do with peaks of latency which make it behave horribly with realtime tasks such as pro audio. It's so bad that it's often perceiveable in desktop usage.
linux-rt does mitigate this considerably, but it's still not very good.
I'm hopeful for genode (with seL4 specifically), haiku, helenos and fuchsia.
And I agree that a small amount of swap can actually reduce the paging that occurs, if you assume that the amount of RAM required is independent of the amount of RAM available. However, as we all know, stuff grows to fill the available space, and if you do configure swap you just delay the inevitable, not prevent it.
Having said that, having swap available means that when memory pressure occurs, you have a more graceful degradation in service, because the first you know about it is when the kernel starts swapping out an idle process to free memory for the new memory hog that you are interacting with. This slows down your interactive session, but not as much as if you have no swap available - in that case, the system suddenly and drastically reduces in performance because it is trying to swap in your interactive process. The more graceful degradation of having some swap available gives you a chance to realise that you are doing more than your computer can cope with, and stop.
As far as I see it, there are three solutions:
1. Disable overcommit. This tends to not play very nicely with things that allocate loads of virtual memory and not use it, like Java. And if you do have a load of processes that actually use all the memory they allocated, then you can still have the same problem occur. The solution to that one is to get the kernel to say no early, before the system actually runs out of RAM.
2. Get the OOM killer to kill things much earlier, before the system starts throwing away clean pages that are actually being used. On my system with 384GB RAM, I have installed earlyoom, and instructed it to trigger the OOM killer if the free RAM falls below 10% (and remember that stuff you are actually using, but happens to be clean, counts as free RAM). This is the easiest and quickest solution right now. If your main objection to this is that you are inviting the system to kill things that might be important, remember that the kernel already does this, and if you don't like it you should use option 1 above (and really hope that all your software handles malloc failure correctly).
3. Introduce a new system in the kernel to mark pages that are actually being regularly used but are clean as "important", and no longer count as free RAM for the purposes of calculating memory pressure. This could either be as a new madvise (but it would be impractical to get all software developers to start using this), or by marking all binary text by default (which would neglect the large read-only databases that some programs hammer), or by some heuristics. This would then trigger the OOM killer (or the allocator to say no, depending on overcommit) when actual free RAM is low.
Things get swapped around and memory is often close to the limit. Linux then becomes unresponsive, and basically stalls. Theoretically it recovers, but that process is so slow that the next stall is already happening.
It is therefore impossible to run large scale Matlab simulations on my Linux machine, while it is no issue in Windows.
As far as I can see, Linux is only usable with enough RAM so that it is guaranteed you never run out. I don't know why this has never been an issue, I guess because it is a Server OS and RAM is planable, or very infrequent?
total used free shared buff/cache available
Mem: 15Gi 6.6Gi 4.0Gi 295Mi 4.8Gi 8.3Gi
Swap: 27Gi 10Gi 17GiNote that Swap is just a piece of compressed memory. I have no real swap and sappiness is set to 100%.
The one trick, SSDs don't like.
I know the solution is to never use more memory than I have RAM, but that's just what happens - and I know there is no way to "solve" the issue and make it magically run. The issue is HOW it is handled... It's weird that other OSs can deal with this, while Linux needs to be restarted.
I think the issue is that only a handful of (Matlab) processes eat up all the RAM, so this "OOM" can not really do anything - there's no use killing off other processes. What should probably happen is to kill Matlab or one of its processes, even though it is in use. I'd be fine with that. Give some out of memory error and kill or suspend the process. At least then we know.
Instead, the system just locks up completely (because the other threads keep trying to push stuff into memory), but is not actually dead, so we don't even catch the issue. Also, because EVERY process is essentially stalled, you can't even kill Matlab yourself. Or suspend it and dump the data, which would be useful. No, you have to hard reset the machine.
What really confuses me is that this kernel was developed when SSDs didn't exist, so how on earth did "The system becomes irrecoverably unresponsive if a single application uses too much RAM" get missed?
I don't know, there are/were several similar issues (very basic situations, frequently encountered by everyone or at least many people) which are/were not fixed for years (we might say decade(s)): that one dealing with memory exhaustion; then right after that, the problem which follows when memory is freed but the system is still unresponsive for several minutes(!); freezing when writing to a USB disk; freezing when something goes wrong on a NFS mount...
I never understood why those common and really important issues were not tackled (or not tackled before many, many years). IMO they were such basic functionalities, which a proper OS is expected to perform reliably as a basis for and before all the rest, it should have been dealt with and granted highest priority.
Servers can be provisioned with way more than enough memory for its use case and can have spares configured to take up the load if it has to be killed.
On the desktop side the issue happens far less if you have more than enough memory. A developer running vim and firefox/chrome on his 32GB of ram machine is vastly less likely to experience this issue than a cheap laptop with 4GB of ram.
I've tried using a VM with overcommit turned off and a modest amount of memory. Among other things, my mail reader, mutt, used more than half the system's memory when looking at my mail archive, so it couldn't fork to exec an editor to write a new mail.
The middle ground is your own config unfortunately - cgroups can limit available memory, but you'll have to set it up by hand.
Linux is completely useless with ram that is almost full in a way that OSX and windows absolutely are not.
One way to make relatively sure that this is always the case is to use a userspace daemon like earlyoom.
Though to stay fair all desktop OS'es behave badly when put under memory pressure, it's just that Linux is an order of magnitude or so worse.
Perhaps the real bug is that Linux distros make it easy to run swapless.
Almost all of them get closed because they get old and no-one is willing to do the necessary work to refine the bug report to dependable reproducibility.
And that's not surprising really; it's hard, time-consuming work which may be obsoleted on the next kernel version.
That said, some have written memory stressors and have been able to crash or stall a machine but it's still a bit hit and miss.
A better one would just make a C program that mallocs more memory than is available. The "open enough tabs so that it crashes part" is like "banana for scale", it is incredible unspecific. You could probably open HackerNews 50x more than espn.com
The VM is basically paging all clean pages in and out constantly as their tasks become runnable. A pretty standard case of thrashing.
Cold pages - not being used
Warm/normal pages - average use pages
Hot pages - being used a lot
1. LRU/histogramming cold RO-code/RO-data pages to be paged out (dropped) or compressed when other pages are hotter and there's memory pressure.
2. Compressing or paging out to disk volatile RW-data when cold and other pages are hotter and there's memory pressure.
3. Pinning in memory (compressed and/or uncompressed) pages that are hot for performance reasons, unlocking them when they become less hot or others become hotter.
4. Having the ability (like on macOS) to signal applications that the system is under severe memory pressure (say "SIGMEM", "SIGLOWMEM" or "SIGOOM"), and to drop volatile data that can regenerated or is ephemeral.
The kernel can evict memory-mappings of executable files which are currently running. When they jump to a part of the executable that is no longer in memory, it can page that part back in from the executable file on disk.
This is pretty cool. But when memory is very low, the kernel will evict practically all user-space executable mappings from memory, and will be reading back in and evicting back out executable file contents on practically every single context switch. It's just trying so hard to squeeze out some space to make its tasks fit in memory and complete successfully.
I think this was the desired behavior of big-iron batch-processing in the 90s? Not sure why it has persisted so long. I'm a big fan of linux and this is my biggest pet-peeve.
If the system is low on memory, a page of program code may be dropped from RAM and re-fetched from disk when it is needed again, i.e. when that section of the code is being executed again.
(this effect is not limited to program code, though)
Poof, no more networking
[1] The big one for me is that I can't run the Atom text editor, Firefox, and a virtual machine all at the same time.
Is this on Linux? I have no problem with that workload on an 8GB Macbook Air (2019). I was apprehensive about going with the 8GB model, but I've not really had any issues.
I wonder if memory compression on OS X helps here.
Have you thought about running browser-based versions? I tend to find that these behave a lot better.
I recommend that you try ZRAM. It has made a difference in a 8GB laptop that I use: https://www.cnx-software.com/2018/05/14/running-out-of-ram-i...
For 8GB I run a 2GB ZRAM swap.
Memory is just a resource. If you can recover from disk space exhaustion, you can recover from memory exhaustion. I think the current standard of memory discipline in the free software world is inadequate and disappointing.
For example, in a Java server application, if one request encounters some buggy code that tries to read past the end of an array, that request will fail, but all others will succeed - this will give a good chance for the system to be usable, and getting a good bug request, with system-generated diagnostics, for the buggy requests.
However, in Go or Rust, the same scenario panics and kills the entire process by default - turning a potentially minor bug in some obscure part of the system into a system-wide crash.
OOM is obviously harder to deal with (e.g. if one request is using too much memory, there's no guarantee that it won't be other requests actually seeing the OOM errors first), so if we don't even want to deal with the easy stuff, how can we hope to deal gracefully with the hard ones?
It's like happens when you read out of bounds in C, except it fails more reliably.
It's certainly an easy to understand solution and is getting the program into a well-known state, but it's also low effort and user-unfriendly. They could have done better, but they would need real exception support for that.
I've set up monitoring that pages me when "dmesg" includes "OOM".
It is possible to do when you have shared servers as well, by letting root log in to another port instead and setting that process to realtime instead.
Thoughts in my head but never got around to do it.
My first Android device was a Nexus One. 512MB of RAM for what is essentially a full Linux system. Able to run a browser and multiple Java apps, all isolated and running their own VM. Task managers often reported near 100% RAM use and things still worked fine.
And my understanding is that optimized things further, but given how overpowered phones are today and how bloated apps are, it is hard to tell.
Used in combination with compressed swap (ZRAM), it greatly alleviates this problem on the (at least GNOME-based) Linux Desktop.
Still, browser really need to do something about the memory problems they're causing. They're Windows 95-level bad at managing their high memory/leak cases - just leave a browser with more than a few dozen tabs open over night. Especially with a tab that does background fetches (e.g. Facebook or Twitter or something with a lot of timer-driven Ajax queries).
I assert that if it weren't for browsers, there'd be no memory problems on modern desktops.
[1] https://lore.kernel.org/lkml/alpine.DEB.2.21.1810221406400.1...
[2] https://lore.kernel.org/lkml/20181024155454.4e63191fbfaa0441...
[3] https://lore.kernel.org/lkml/20181023055655.GM18839@dhcp22.s...
both use the recently added pressure stall information (PSI)[0] infrastructure in the kernel to determine when the system is overloaded.
I just arrived at this thread after my entire system stalling completely at yet another low memory situation.
Let's just say I'm extrememly grateful to discover some of these userspace early OOM solutions in this thread.
Sometimes instructing the kernel to clear its caches helps: `echo 1 | sudo tee /proc/sys/vm/drop_caches` [1]
[1]: https://serverfault.com/questions/696156/kswapd-often-uses-1...
In that case, it cannot be moving things the swap since it was alreay full. What did probably happen is that since the swap was already full and not enough it started to trash the system by paging out the executable files in the memory. Ironically enough you'd probably have much smaller risk of trashing had the system have more swap space.
Without swap it just kills processed until there is enough memory, which is what you would have done anyway!
I think the main annoyance with Linux here is that in Windows you get to choose what to kill, whereas in Linux it can't really communicate with you (because the kernel doesn't know about such modern things as GUIs) so it had to pick more or less randomly.
for P in /proc/[0-9]*; do echo $(cat $P/oom_score) $(cat $P/comm); done | sort -n | tail -5
For me right now that shows 4 "Web Content" processes (firefox tabs) and a firefox-esr. That seems to check out.In normal conditions this frees up memory for more useful data and helps you avoid getting to perverse conditions.
Then again, this "allocation will never fail" mentality has also lead to applications being written with such an assumption, and when allocations do fail, they crash. (Arguably, that's better than thrashing the rest of the system.) I don't know if the modern browsers will actually stop letting you open new tabs and just give an "out of memory" error instead of crashing, but that's how most Windows programs are usually written --- without the assumption that allocations can never fail, because on Windows, they can.
It's not data swap, it's executables. Linux knows it can reread the executable from disk if needed, so it uses those memory pages from other things, and reads them in when needed.
There nothing really special about anonymous memory except that it's backed by swap instead of by a named file on some filesystem. On a system with swap disabled, the backing store is still conceptually swap, but since the swap doesn't exist, pages backed by that imaginary swap can't be evicted from RAM. Pages backed by other backing stores can certainly be evicted from RAM, however, and that's how you "swap" on a swapless system.
Note that executable code is almost all mapped so that it can "swap" in this way.
What I've always thought is that there should be a working set size limit on a process which includes the buffer cache somehow. The idea is that the process may not use more RAM than this size- if it exceeds it, it must either fail or swap out its own pages, not those from any other process. This would fix the problem for tar- it only needs a tiny amount of memory.
I think the situation is very similar with the web-browser example. The browser should not be allowed to force all unrelated data to be paged out.
You can control this with cgroups. Plug a process into a separate cgroup and set the `memory.limit_in_bytes` knob to whatever your heart desires.
I use it to limit the qBittorrent's memory usage on my machine. `firejail` is very convenient for doing this. If I don't set a limit (30% RAM in my case), it eats up all the memory with a uselessly large file cache, which does not improve upload speeds at all.
https://access.redhat.com/documentation/en-us/red_hat_enterp...
Also if someone opens up an application that grabs huge chunks of RAM but leave a lot of it idle, and turn swap off completely, they should not be surprised. I don't know why people see this as a bug, but perhaps I've just been spending time in the UNIX family tree for too many decades.
12 years ago, before I got a mac, I had a windows XP laptop with enough memory (8GB) to disable the swap file (which world+dog will insist is a very stupid thing to do). This was great and vastly extended the useful life of my laptop. Alt+tabs were instant and I could run e.g. JVM applications with sane heap settings as well as a browser, office stuff and a few small things I needed with zero issues. On the rare occasion that something did run out of memory, it died or I killed it. Laptop disks were stupendously slow at the time; any form of swapping on a slow laptop disk is extremely disruptive. SSDs are much better but there too it tends to be mostly redundant.
IMHO most forms of swapping are highly undesirable on both servers and end user hardware. Swapping to free up memory for file caching is simply unacceptable when you can instead just evict cache pages. If you don't have enough memory left to cache effectively, that just means things like memory mapped files will get a lot slower. If something allocates more memory than you have just kill it.
A small swap space (~ 1G on a 64+G ram server) is a reasonable backstop against a slow memory leak. Assuming you don't have filesystem pages evicting anonymous pages, swap use is a clear indicator of too much memory use and points you in the direction of something to fix; and gives you a little bit of time to fix it on the running system. As long as swap is very small relative to ram, it's not going to enable thrashing -- a big leak or burst in use isn't going to fit in swap and you're going to be dead anyway.
Understatement of the year. If you can't cache your memory mapped executable pages, your system will be just as slow as in the nightmare swap scenario.
Still open.
If the program tries to use more RAM it'll then just die, and not drag down the whole system with it. Works fine, but I really shouldn't have to do this.
Alternative link, just in case: https://lore.kernel.org/lkml/d9802b6a-949b-b327-c4a6-3dbca48...
To me, this is the Linux kernel's biggest weakness against Windows. Most other gripes about Linux (poor power management, poor driver support, etc) belong outside the kernels domain, but this one is a glaring win for Windows over Linux.
Besides which, I'm looking at https://github.com/torvalds/linux/blob/master/mm/oom_kill.c and I genuinely can't see where root processes get privileged. I can see a reference to a `root bonus` in a comment, but other than that... maybe my kernel source reading skills are just too rusty.
Something needs to be done here for real, otherwise Linux is a nice software
I can't figure it out from the answers. I think with root privileges it should just be possible to say "GUI has higher priority" etc... Then when there is a memory issue you kill some low-priority processes to get the memory back.
But whatever any sane person considers part of the operating system because it is the bare-minimum of what is required to do stuff (filesystem, gui, ...) needs to have priority and always be fast. This can be defined by the distribution using the startup privileges.
So, why is this so difficult?
idoubtit 2 hours ago | unvote [-]
Point 3 is wrong. OOM killing is not random. Each process is given a score according to its memory usage, and the highest score is chosen by the kernel. The way to mark priority in killing is to adjust this score through /proc. All of this is documented in `man 5 proc` from `/proc/[pid]/oom_adj` to `/proc/[pid]/oom_score_adj`. http://man7.org/linux/man-pages/man5/proc.5.html
I'm not sure if it'll exactly accomplish what you want though.
Are you recommending a swap space be created automatically on behalf of the user?
One could also use compressed memory (`zram`).
Nonetheless, I'm surprised someone is calling this a bug. Let's face it, Linux is just not a desktop operating system. It's a server operating system, and it expects that it will be professionally administered and tightly controlled to prevent OOM situations. That OOM situations occur on servers too is beside the point. There are reasons for the linux memory system to work as it does, reasons Linus will yell at you about if you complain.