Malloc Never Fails (2012)
scvalex.net
scvalex.net
> To clarify, the surprising behaviour malloc has does not mean we should ignore its return value. We just need to be careful because malloc returning successfully does not always mean that we can use the requested memory.
It's worth pointing out that if you turn overcommit off (https://www.kernel.org/doc/Documentation/vm/overcommit-accou...) you will, in fact, get this guarantee
> if you turn overcommit off
The author said ‘does not always mean’ so giving one condition in which it does doesn’t prove them wrong does it?
If this were an option on a malloc call, any such call would have to reserve the allocated memory at that point, reducing the usefulness of overcommit for other processes. This would set up a 'tragedy of the commons' scenario, where every application developer defensively uses this feature because other applications are using it.
I don't know quite what 'enough' means, but this whole business is already a rat's nest of heuristics, so one more should fit in nicely.
I discovered this because of differences in how jemalloc and my system allocator (from glibc) allocated memory. Namely, jemalloc would pass the MAP_NORESERVE option, which meant my program behaved as if overcommit was always enabled. Thus, an out-of-memory condition meant that it was subject to the OOM killer. But when I used my system's allocator, the MAP_NORESERVE option was not passed, which led to the heuristic detecting an out-of-memory condition and causing the malloc call to fail.
For more info, see `man 5 proc` and/or my exploration here: https://github.com/BurntSushi/ripgrep/issues/993#issuecommen...
(It's all still a bit wishy washy, so I think your point still stands. But I figured I'd chime in with some bits that don't appear to be widely known, and which led to some surprising differences in behavior based on which allocator one chooses.)
The issue with this is that there are other processes on the system. If I start a program that used your idea for allocation, and it used 90% of my memory, a small growth in my memory usage might mean I run out. But because it's guaranteed not to OOM, another program gets killed rather than the one using 90% of the memory. If you want no OOM kills, you've got to disable overcommit globally.
For example, fork() will have to ensure that there is enough memory for a complete copy of the running process. If you have a large process, e.g. using 4GB of a 8GB machine, then fork() won't be able to run, even if you just want to fork and run a tiny program.
With over-commit turned on, the fork() would work (because of copy-on-write) and the program could then happily exec() the new program.
The work-around is to allocate huge amounts of swap space so that the OS can be confident that it can reserve all the potentially required memory from fork(), even if it never normally has to use all that memory.
But many Linux applications assume the overcommit behavior. Its far easier to just "go with the flow" or "When in Rome...". Overcommit is the programming culture of Linux and should be assumed when writing Linux apps.
intentionally. you don't need to point out the "error". it's a rhetorical device.
Linux overcommit, and developer mindset is a detriment to software quality and portability.
It's very Linux-centric and presumes a certain config+usage pattern. Not true on Windows. Not quite true on macOS. Not true in WASM. Definitely not true on embedded platforms.
1. Rust the language knows nothing about allocation. If you care about this behavior, it mostly limits the code of others' that you can use, but you can always write your own versions of things that respect fallible allocations.
2. Rust's standard library assumes memory is infallible. This is partially because it's a good default, and partially because our allocator API was not ready yet.
3. We've been working on the allocator API.
4. We have a rough plan for parameterizing data structures over allocators.
5. If this topic is of interest to you, https://github.com/rust-lang/wg-allocators is where to get involved.
And Intel has a spec out for a PML5 page table, giving you 57 total virtual address bits.
I really should have coffee before HN
Linux's Overcommit behavior is non-obvious to many programmers. Its one of those issues that very few programmers I've come across in the workplace understand properly.
This blogpost properly understands the issues associated with Overcommit, and have done some preliminary investigations that describe the behavior. Its a really good blogpost.
The general point of the blogpost is that "Malloc Fails due to address space exhaustion more often than actual memory-exhaustion". Because actual memory exhaustion causes OOM killer code to be run... and OOM killer is a non-obvious case of the Linux kernel.
More generally, the article does nothing to educate the public or move the "debate" about overcommit in any sort of useful direction.
Since glibc malloc is built on top of mmap and brk, I think your distinction is mostly academic. For any programmer using Linux and glibc... malloc's failure mode IS mmap and brk failure modes.
Of course none of this applies to arm/aarch64.
Set VM Overcommit to zero on an embedded system with no swap (a Raspberry Pi will do nicely).
Write a C program that malloc()'s all the RAM.
Watch malloc start to fail when you hit the RAM limit and the kernel has dumped all the I/O cache it can.
The larger point is that very few users of malloc understand its semantics, and in fact you can't know exactly how malloc will behave without knowing things about the runtime configuration of the system (as opposed to the hardware availability as many people like to think the simple case is).
Not that you can do much when malloc fails...
Tell the user they can’t do that operation? Use on-disk storage instead? Re-use memory you already have (such as evicting a cache and taking the memory it already had allocated)? Abandon what you were doing if it was only an optimisation and wasn’t essential (such as allocating memory as part of a speculative execution)? Run a garbage collector and try again?
Lots of options available in some situations.
Once 32bit hardware-based virtual memory entered the picture, every single application would just assume you have endless RAM - memory checks are reserved for "common cases" like trying to create a 100000x100000 image in a (32bit) image editor.
The reason for all that is simple: out of memory situations are stupidly and increasingly rare and in most cases where they can happen there isn't much you can do (e.g. what would you do if you run out of memory while making the fourth button in a toolbar?) and really in the 99% of the cases there might only be five people in the entire universe that will encounter such a case (two of them will tweet about it though and amass a lot of "lol, those garbage developers" retweets) so littering your codebase to keep those five people (and their Twitter followers) happy is not worth the effort. I mean, are you really going to put a "run out of memory" check after every toolbar button allocation? And what are you going to do if that fails? What can you do?
AFAIK some modern languages nowadays even assume memory allocations wont fail (and if they do they just terminate).
Rollback what I’d created so far, evict caches, ask malloc to trim, try again, if that failed roll back again and ask the user to reduce the volume of application data open by closing views or documents or whatever before they try again.
But yeah it’s a lot of engineering.
Silly engineering like what you suggest is a waste of time. Non-critical applications that are allowed to crash (every normal desktop app) would end up overengineered and unmaintainable. Embedded systems that must give guarantees about their reliability simplify system design drastically. One of the things going out the window first is dynamic memory allocation. It has too many failure modes.
You just failed to create a toolbar button, how are you going to ask a user do something if you already failed to create a tiny UI element?
So the typical behavior you saw was a combined effect of physical constraints and lack of skill or care.
Exiting cleanly would mean trapping SIGBUS/SIGSEGV (can't remember which one is invoked) and being able to roll-back from just about any point in your code or libraries)
You need to ensure overcommit is disabled, but because that's a system-wide setting, your program can't easily force that on.
I have seen code hit an allocation failure, bubble the error up the stack, which causes various pieces of code along the stack to free their heap allocations, all the while the process is able to log what is going and keep running.
I have even seen this happen when the allocation that failed was tiny, less than 100 bytes.
I don’t know, maybe this is old school, but yes you should put a check on every toolbar allocation for two reason: one you can gracefully report why your app isn’t doing it’s job and two it is a very very slippery slope when you start ignoring those errors as a matter of practice
What if you cannot do that report because the reporting itself needs memory that can fail?
> but yes you should put a check on every toolbar allocation for two reason: one you can gracefully report why your app isn’t doing it’s job
This is a reason to put checks in every toolbar allocation, but on the other hand there is also a reason to not do that: you are trying to catch a case that will only happen at a 0.00001% of the time yet to do that you are making the codebase more complex which will affect working with it 100% of the time.
Unless you are working in something like a medical or nuclear device, it is not worth to bother with such things. Even then when it comes to high risk programs relying on the programmer doing things right outside the core functionality is probably (i haven't worked on such projects myself so i'm guessing here) not a good idea and instead you should compartmentalize (e.g. the core functionality is running as a separate process from the UI and if the UI crashes, a watchdog restarts the UI process). But that is my guess here, i mainly have common end user/consumer applications in mind.
Where "something like a medical or nuclear device" implies a piece of software that is expected to work correctly.
In other words, you should always ensure that failures resulting from malloc(3) returning an error (and all other errors) are suitably contained.
My implication was software that if it fails it will kill people.
Otherwise if a program crashes because of a situation happens once every 100000000 runs, not only is worthless to worry about it, it actually is preferable to not do that as to keep the codebase clean and hence easier to maintain for bugs that actually do affect people.
Allocate and reserve error-reporting memory early. For example, your ErrorReporter instance is initialized at start-up, and doesn't call malloc() when you make a report.
Allocate a big block of memory at startup. Free it just before doing the reporting.
If the out-of-memory was caused by something simple like accidentally loading 2G of data because of some odd data in an network request or a user selected file, and it rolls back to the start of the UI action or the incoming network request, the application may still work fine.
(I usually connected other out-of-resource conditions such as pipe/socket/open failures to std::bad_alloc too. They're pretty comparable in effect and give a bigger chance of actually exercising those exception paths)
If you have a system that handles independent requests, it might be more robust to model each request as its own process, and then simply allow those processes to crash if they find themselves in an unexpected condition (and handle process crashes).
In eg an application server, requests themselves may be independent, but still share a lot of cached data - and as long as locking overhead doesn't overwhelm you, a single process is still the fastest way to share data between (worker) threads.
If your threads need to share interleaved modifications to some common state then that approach isn't viable. But in that case it will be very hard to "roll back"/"skip over" a failed request, as it's difficult to be confident that you haven't corrupted that shared data in a way that will cause the same problem for future requests.
I'm not sure why skipping over failed requests is hard. VMs associated with the request are aborted, the client gets a 500/503 and can retry later. Incomplete data doesn't get added to the caches at all, so no corruption of shared data.
memset(malloc(size), 0, 1)
The program will crash if malloc returns NULL or the memory is not writable.This is probably outdated, given that it was written in 2009.
(Yes, I realize Linux is the kernel, etc.)
> The malloc function allocates space for an object whose size is specified by size and whose value is indeterminate.
> The malloc function returns either a null pointer or a pointer to the allocated space.
I think it's certainly debatable whether overcommitting while lazily allocating counts as actually allocating.
> The malloc function returns either a null pointer or a pointer to the allocated space.
Uhm, either malloc returns a pointer to the allocated space, or it returns null. I don't see what's so complicated about this. There's no provision for "return a non-null pointer to unallocated space".
Section 7.22.3.1:
The order and contiguity of storage allocated by successive calls to the aligned_alloc, calloc, malloc, and realloc functions is unspecified. The pointer returned if the allocation succeeds is suitably aligned so that it may be assigned to a pointer to any type of object with a fundamental alignment requirement and then used to access such an object or an array of such objects in the space allocated (until the space is explicitly deallocated). The lifetime of an allocated object extends from the allocation until the deallocation. Each such allocation shall yield a pointer to an object disjoint from any other object. The pointer returned points to the start (lowest byte address) of the allocated space. If the space cannot be allocated, a null pointer is returned. If the size of the space requested is zero, the behavior is implementation-defined: either a null pointer is returned, or the behavior is as if the size were some nonzero value, except that the returned pointer shall not be used to access an object.
Section 7.22.3.4:
The malloc function allocates space for an object whose size is specified by size and whose value is indeterminate.
> The pointer returned if the allocation succeeds is suitably aligned so that [...]
> If the space cannot be allocated, a null pointer is returned.
If you use mmap and "from your perspective" that's considered success then it's your perspective that's faulty. The standard is literally telling you right here^ there are 2 possibilities: either you allocate the space and return a pointer to the space, or you don't and you return null. There is no third option of "space cannot be allocated but you return non-null anyway". That's quite literally the end of the story.
Exactly! That's why once virtual memory is allocated, malloc() is allowed to consider the operation successful. The standard does not care at all whether it is virtual memory allocation or physical memory allocation. It is completely unspecified in the standard what sort of memory must be allocated. So no spec in the standard is being violated by returning non-null pointer for virtual memory allocation.
So once again, can you cite the exact section number from the standard that you think is being violated here?
Ironically, it's a memory model that greatly reduces the performance of modern systems, and als impacts its safety.
We can't expect a kernel to be aware of all language specifications.
Yes, I know that in this case C is both the program's and the Kernel's language in this example but even then. Languages go through iterations (versions) and a Kernel can't be required to obey all languages which might run under it.
Linux (the kernel) doesn't have malloc. It's part of the C library, which is a completely separate project. What Linux implements is brk and mmap.
But even then I'm not convinced that anything here is non-standard, at worse maybe we're in a bit of a grey area. As long as the kernel maintains its smoke and mirrors whether and how it allocates memory is irrelevant from the point of view of the standard. The C language has a rather simplistic memory model, it doesn't impose a lot on the implementation.
Now the problem occurs when the program attempts to access virtually-allocated memory and the kernel realizes that it can't find any physical memory to map it to. In this situation several things can happen but in general the process will be killed. Is it against the standard for the OS to kill a program for arbitrary reasons? I can't imagine why. It could also freeze the program, waiting for more memory to become available. Again, not against the standard as far as I can tell. Or maybe kill some other program to free memory.
If you have some specific part of the C standard in mind please do tell, I always find these language lawyering arguments interesting, somehow.
See here: https://news.ycombinator.com/item?id=20145604
Also note the POSIX standard:
Upon successful completion with size not equal to 0, malloc() shall return a pointer to the allocated space. If size is 0, either a null pointer or a unique pointer that can be successfully passed to free() shall be returned. Otherwise, it shall return a null pointer and set errno to indicate the error.
In neither case is there any provision for returning a non-null pointer to anything other than an allocated block of memory of at least the given size.
We're not, and we never were, debating the situation where dereferencing the non-null pointer returned by malloc succeeds but takes long time due to your Amazon order. We've been talking about the situation where it fails. malloc is not allowed to return a non-null pointer to a memory block that cannot be written to. Linux does it anyway, and in doing so blatantly violates the standard.
char *b = malloc(2);
if (b == NULL) {
return 0;
}
b[0] = 'A';
b[1] = '\0';
printf("%s\n", b);
free(b);
The standard tells me that if the malloc succeeds then the following code, if allowed to run, will display "A" on stdout. The C standard cannot and does not guarantee that a C program can't be interrupted however. For the sake of the argument we could imagine a kernel that instead of killing the program freezes it indefinitely on disk waiting for RAM to be available. It's functionally the same thing. As long as the kernel doesn't let code run with broken invariants it's fine. This is completely outside of the scope of a language standard to define.Or, to try one last time from a different direction, if you consider that the C standard mandates that accessing memory returned successfully by malloc has to be successful and I happen to press ^C when that happens in a program, should the kernel refuse to kill the program? This is obviously absurd, but it's effectively the same thing: the kernel reacts to some external state and decides to terminate the program.
P.S. I don't think indefinite hold is "functionally the same thing" as termination. A caller system(), for one, would need to return in one case, but not the other.
See 5.2.4.1 in http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1548.pdf
However, it's true that not any old instance of this malloc problem demonstrates such a nonconformance. If it happens in a large program that has allocated gobs of memory, then no.
Basically if the system is low on memory that it can no longer support the execution of a small C program with modest memory use, then it becomes nonconforming.
However, the mere property that memory can be doled out by malloc which might later not be used doesn't make it ipso facto nonconforming.
Moreover, a system with any kind of memory management (including management that earnestly reports null for "out of memory") can be come a nonconforming C implementation if it is low on memory.
Both the hardware burning after memory allocation and lack of availability of physical memory after memory allocation are outside the scope of the standard. The standard says nothing about them. A C program can fail in these scenarios without violating the standard.
Maybe it's more accurate to say that Linux is a kernel that assumes many features and semantics of the C runtime. But it certainly seems more deeply intertwined than you're willing to address here.
We may be able to argue that it doesn't. ISO C says in 1. Scope, this:
This International Standard does not specify
[...]
— the size or complexity of a program and its data that will exceed the capacity of any specific data-processing system or the capacity of a particular processor.
It seems that these weasel words have an interpretation that can be bent around overcommitted memory allocation.
For an implementation to be deemed conforming, it just has to be demonstrated to successfully translate and execute one program that tests each of the minimum implementation limits.
See 5.2.4.1 in http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1548.pdf
Only if no such program can be found can we then conclude that the implementation is nonconforming (and if the reason for not finding such a program is this overcommit issue, then we can blame that issue).
Your Linux system is indeed nonconforming if it is so low on memory that, for instance, no C program can allocate a 65535 byte object (that can be actually initialized, and used: a real object). Basically the memory situation has to be so severe that it takes the implementation below the minimum limits, whereby we can clearly demonstrate nonconformance.