Dealing with Out-of-Memory Conditions in Rust
crowdstrike.com
crowdstrike.com
So if your aim is to build a reliable system, it is much easier to plan to never get there in the first place.
Alternatively, make your application restart after OOM.
I would actually prefer the application to just stop immediately without unwinding anything. It makes it much clearer as to what possible states the application could have gotten itself to.
Hopefully you already have designed the application to move atomically between known states and have mechanism to handle any operation getting interrupted.
If you did it right, handling OOM by having the application drop dead should be viewed as just exercising these mechanisms.
How about an image viewer that tries to open too large an image, should it just crash when it OOMs? I would much prefer an error dialog and the program continues to run.
As a user I am ok with this. I don't let my memory end and would prefer that the application worked reliably while there is still memory.
What I don't like is applications trying to pretend nothing happened but then doing some strange shit.
I have seen, for example, applications using 100% cpu after they hit OOM. Yes they survived. No, it is not user friendly.
This is a very reliable and easy for the user and other developers to understand.
try
{
image = LoadImage(path);
}
catch(OutOfMemoryException e)
{
Msgbox("Cannot load image it is too large for available memory");
}
That is way more friendly than a program crash and allow the user to try again perhaps with a smaller version of the image because they accidentally picked the high res version or something.Either way you may get 100% CPU as the program crashes and memory gets reclaimed, or the program continues to run and the garbage collector reclaims.
After that if there is still not enough memory for the simple message box then you have an uncaught exception and the program crashes and you're just back to the no catch approach, nothing lost. Most likely though there will be enough memory for the message and you have a much more friendly result.
I have used this approach before and it is way way more friendly then a crash with users potentially losing work because they accidentally picked the wrong file.
That's exactly the point. Once memory is exhausted, you can't take any action reliably.
You don't build reliable applications by making mechanisms like: "ok, if memory ends lets design it to show a box to the user, we have 50% chance this succeeds".
It would be better to wrap your application in a script that detects when the application quit and only then shows a message to the user.
People designing things like you drive me crazy. They come up with a huge number of contingencies that just don't work in practice when push comes to shove.
Stuck in a loop trying to show a widget, using 100% of CPU and preventing me to do any action?
I prefer simpler mechanisms that work reliably.
The trouble is that you may not know beforehand what exactly is going to be needed. And you might need maybe a library call and the library does dynamic allocation in it and you either have no idea about it (until you find out the hard way) or no way to help it.
So in the end maybe you can take some extremely simple action like writing something to log or show a widget, but that's about it.
Some PLs come with test suites that give you a failing allocator, so you can even easily test to make sure this looping condition resolves sanely.
Nothing is 100% reliable, thats not realistic, and in this kind of situation its is not a 50/50 shot, its more like 10000/1 that you will be able to show a message, that should be obvious.
>It would be better to wrap your application in a script that detects when the application quit and only then shows a message to the user.
Thats simpler then wrapping the specific function in a "script" that shows message to the user but allows the program to continue to function in almost all situations?
>People designing things like you drive me crazy. They come up with a huge number of contingencies that just don't work in practice when push comes to shove.
People like you who do not value you the user experience over pure code drive me crazy. These things actually do work in practice, I guarantee you any sufficiently complex GUI program will have code like this to try and gracefully handle as many contingencies as possible before simply crashing. Do you think your browser should just crash losing other tabs if a web site loads too much data? Does Photoshop just crash if it runs out of memory on an operation losing all your work?
>Stuck in a loop trying to show a widget, using 100% of CPU and preventing me to do any action?
Its no more stuck than your process crashing and the kernel is reclaiming memory, they both take similar cpu time, one however results in a message that informs you of what happens and leaves you with a running program, the other tells you nothing and your program is now gone.
My comments were directed at people who are interested in building reliable systems.
If you straight assume it is not possible to build reliable systems, you are missing a lot.
Memory allocation can fail, network connections are not reliable , opening a file may fail, writing a file may fail and so on. If your program simply crashed because some operation failed it would only be reliable at crashing.
By definition, you are not functioning normally if you're OOM.
If you write a large file to disk filling it to capacity, failing then delete the file, now you have free space again and everything is normal.
You've never used btrfs, have you? :D
This is nontrivial in several PLs.
As the article notes, this is all moot if you are running on Linux though. The allocation will always succeed but if it's too big the kernel will start killing processes, quite possibly including your image viewer.
This page explains more on this: https://sqlite.org/limits.html
It's takes very little memory to show the user an error message it will almost always succeed even if the operation that triggered failed due to OOM.
Handling this is hard in desktop applications. For servers, you can have known workloads with better limits on each process.
Under that scenario, the antivirus sensor should be able to take some action (log, likely) that a malloc failed, and possibly even try to recover memory and identify the risk.
But beyond that narrow use case, you’re right.
The application saved state to internal database, and even if you cycled power it would just come back to the same screen and same state it was before power cycle.
This wasn't to deal with memory problems (in fact, it had no dynamic memory allocation at all) but rather to deal with some crappy external devices and users that would frequently power cycle things if it didn't progress for more than couple of seconds.
Why everybody assumes every developer doesn't know what they are doing just because most of them don't?
On the other hand, if you are likely to be identified as the culprit, I think the best you can hope for is getting some cleanup/reporting in before you're kill-9'd.
You might be happy to learn that the OOM killer can be configured[1] to specifically protect certain processes. If the entire point of a machine is to run a single process, then you should definitely use that feature.
Behold, CS is as bad as governments with all their tax rules
This is more useful than OOM'd programs silently disappearing, but isn't foolproof (that's why I said "increase the probability" rather than "guarantee"): if the OOM killer gets invoked before your telemetry, if you've forked, or if whatever your language does in order to even detect an allocator error is itself memory-costly (looking at you, Python--why on earth would you allocate in order to construct a MemoryError object?!), you will still go down hard.
Other than going back in time and reversing Linux's original sin (allocation success is a lie), swap is the "solution" to these situations, but for many people that cure is worse than the disease.
I often wish there was a generally-available way to map memory to files (rather than mmaping files to memory) selectively in my programs. For cases like this--handling OOM conditions and doing cleanup/reporting before turning out the lights, when the OOM condition was caused either by code in my program outside of my control or other programs on the system outside of my control--having a "all allocations after this instant should occur in a file on disk, not memory" bit to flip would be nice. I don't need my cleanup to be fast; if this is happening I'm already on a sad enough path that I'm happy to trade shutdown/crash performance for fidelity of diagnostic data and increased probability of successful orderly cleanup.
Of course, not being an OS developer, I assume this is impossible for reasons outside of my understanding; the few complicating factors I can think of (copy-on-write pages, for example) are hard enough, and I'm sure there are others.
This behavior is the reason Java had such a bad time with sysops.
Just make use of proper libraries that will invalidate caches if needed, or simply refute to spawn new threads/actors/processes.
I’d rather have some application halt/pause or simply crash if it ran out of memory.
This is absolutely not why Java had/has a bad rap memory wise--that has more to do with a combination of code that allocates irresponsibly on the happy (not memory-error-anticipating) path, and the JVM's preference to preallocate as much as possible for every purpose, not just OOM handling.
> Just make use of proper libraries that will invalidate caches if needed
What if I've dumped every cache I can and I'm still getting allocation (or spawn, or whatever) errors because the system is out of memory? It's useful to have a contingency in place to turn out the lights room-by-room rather than cutting power to the whole building, as it were.
• There's no problem of "untested error paths". Rust has automatic Drop. Cleanup on function exit is written by the compiler, and thus quite dependable. Drop is regularly exercised on normal function exits too.
• Overwhelming majority of Drop implementations only free data, and don't need to allocate anything. Rust is explicit about allocations, so it's quite feasible to avoid them where necessary (e.g. if you use a fat error type that collects backtraces, that can bite on OOM. But you can use enums for errors, and they are plain integers).
• Probability of hitting OOM is linearly proportional to allocation size, so your program is most likely to hit OOM when allocating its largest buffers. This leaves a lot of RAM left to recover from that. It's basically impossible to exhaust memory up to the last byte.
Crashing may be fine for tiny embedded software or small single-threaded utilities. However, crashing of servers is expensive, and may not even work. When you crash, you kill all the requests that were in progress. This wastes work already done by other threads, and you're not making progress. When clients retry, you're likely to get the same set of requests that lead to OOM in the first place, so you may end up crashing and restarting in a loop forever. OTOH if you detect OOM and reject one offending request, you can keep making progress with other smaller requests.
Think in terms of trying to save the file you are working on before the application exits. Or notify cluster map that the node is not going to be available to process requests.
How does compiler factor in it?
Each function can fail, and when it fails, you propagate the error up to a point where the whole failing task can be gracefully cancelled.
e.g. if user invokes a "Print" command, and it runs out of memory, then instead of immediately crashing and burning, you can try to report "Sorry, print failed". That too will be handled fallibly, so if `message("Sorry…")?` fails, then you proceed to plan B, which may be log, then save and quit. If these fail, then finally crash and burn. But chances are that maybe print preview needed to allocate lots of memory, and other functions don't need as much, so your program will survive just by aborting that one operation.
Basically there's only one kind of software in this class: DBMS software. DBMSes want to be able to try to process ridiculous queries if users ask them to; and then fail in a way that only affects the processing of that query, rather than the stability of the DBMS as a whole. And they also mostly can't afford the overhead of pre-calculating just how ridiculous a query will be, before attempting it; because that calculation often requires effectively 90% of the work involved in actually running the query.
For every other type of software, letting the OS handle the OOM (by killing your process) — and setting up your higher-level inter-process / inter-node architecture to be resilient to that — is the sensible approach.
Additionally I have personal experience with network appliances (firewalls/deep packet inspection/etc). If such a device runs out of memory, it degrades gracefully (unless there's a bug) and starts shedding connections rapidly to avoid running out of memory. Such devices can't just restart if they run out of memory as that would be a network outage and can often be exploited to cause a more serious denial of service attack.
Rust is a system's language. Handing out of memory conditions is par for the course for systems programming.
As an example, ordinarily on Rust this does what it looks like it does:
let mut lyrics = "Never was".to_owned();
lyrics += " a cornflake girl";
With cfg(no_global_oom_handling) this can't work, most importantly because the second line is an arbitrary concatenation and therefore allocates, but there is no possible way to signal if it fails.How hard would it be for a text editor to highlight all lines that do or might allocate? Can they be known statically?
But I feel like a static analyzer could clearly identify “this line can allocate” just by knowing what language features can allocate. And for third party library methods “does this function have anything inside of it that could allocate?”
fn maybe_alloc() -> Vec<usize> {
if collatz_conjecture_is_true() {
vec![42]
} else {
vec![]
}
}
fn main() {
let v = maybe_alloc(); // Does this allocate?
}You can't know for sure if all of those places are reachable in a running program, but that's not necessary to solve this problem.
Giving this up doesn’t seem very impactful and I’d be happy to do so to enable actually useful features.
I mean, they did warn you: No Global OOM Handling. So, if this mustn't fail, and it might not succeed, therefore it has to be eliminated, a String under cfg(no_global_oom_handling) can't push_str() and it can't reserve() and it can't do a lot of things. Their definitions are conditionally removed from the type.
I didn't check, it's possible "something".to_owned() is similarly forbidden, seems like that would allocate memory.
Basically if you live in a world where allocating memory for strings, vectors, and other growable types feels extravagant, cfg(no_global_oom_handling) is for you, and if not then maybe you should re-evaluate why you are worried about allocation failures when you are wasting precious heap memory on such data structures.
Thanks for the clarification. I had totally missed the point.
Zig has a similar approach that is pretty cool. I don't know of any other language that let's you handle it like this https://ziglang.org/documentation/0.8.0/#Heap-Allocation-Fai...
I would like to read about their successful test of their C++ pieces using the same (random fail small allocations) strategy. My impression from Herb Sutter is that on the popular C++ Standard Library implementations (for MSVC, Clang, GCC) this actually doesn't work, but of course Herb has a reason to say that - so it'd be interesting to hear from somebody who has a reason to believe the opposite.
This might be feasible with some limited-in-features apps, like an audio/video player, an IM client, etc...
Is there a way to disable it for a single program or cgroup, to enable it to deal with out-of-memory conditions? Maybe changing/hooking the standard library?
mlockall(2) seems overkill, since it will also force all code and mapped files to be resident.
I wouldn't be surprised if the popular allocators - jemalloc/tcmalloc/scudo or even the defaults in glibc/musl support a setting like this.
Maybe what I'm asking for makes no sense though, because even if my process handles out of memory errors gracefully, it might still get OOM-killed when another process allocates some more.
The entire problem is, in a way, unsolveable with current OS APIs. AFAIK there is no preexisting, good, actually usable, and universal memory usage meter. Some things work for a lot of cases, but I don't think there's anything that would work universally.
Coincidentally, in my eyes the best way to handle memory pressure in applications would be proactively rather than reactively. Sometimes you can unload more things, like pages in a document that aren't being edited right now. And if you need to fail, you can fail on good boundaries (e.g. refusing entire requests in server-like things) rather than in the middle of something that you need tons of work to unwind correctly.
(Slightly related: pool memory allocators, e.g. APR's https://apr.apache.org/docs/apr/trunk/group__apr__pools.html )
But at the same time it's critical for kernel and embedded (µC, the "too small for Linux" kind) code.
Isn't it only Linux that does overcommit by default?
I use https://lib.rs/cap which self-imposes memory limit on the process. Setting that limit below cgroup limit allows programs to actually handle OOM before they get killed by the OS (if only Rust's libstd wasn't so eager to self-abort anyway).
How do you choose that limit?
For desktop applications it's tough. It may be just an arbitrarily high amount you don't expect to hit during normal operation. If you need to work with variable-size data, then it could be `size_of_file_being_opened * x` if you can predict the `x`.