Once upon a time, memory allocators made sense
github.com
github.com
This has two effects:
(a) malloc() never returns NULL. It always returns a valid address, even though your system may have out of memory.
(b) by the time the kernel finds out that it's run out of physical pages, your process is already trying to use that memory... which means there's no way for your process to cope gracefully. You have to trust the kernel to do the right thing (either to scavenge a page from elsewhere, or to kill a process to make space). If you're very lucky, it'll send you a SIGSEGV...
It's really annoying, your application may get killed without any way to react, even to just print a "Low memory" error.
You can run out of both physical memory and swap space. Or your system is swapping is so heavily that, for all practical purposes, the system becomes unusable.
I should probably add this to my list of non-security reasons (or an angle at least) for using Nizza-style architecture [1] where critical functionality is taken out of Linux part to run on microkernel or small runtime. Maybe even potential to go further and have such apps that monitor for this sort of thing with a signal to apps that works across platforms. Not just this design flaw, but potentially others where app doesn't have necessary visibility. What you think?
This is only true on Linux, for example, FreeBSD by default has partial overcommit and WILL start returning null after overcommitting some percentage of memory (I think 30% or so by default)
You can also enable behaviour like you describe by setting overcommit_memory=2 and some ratio overcommit_ratio (default 50) or exact number of kbytes overcommit_kbytes
Source: https://www.kernel.org/doc/Documentation/vm/overcommit-accou... https://www.kernel.org/doc/Documentation/sysctl/vm.txt
Unless of course you run out of virtual address space.
This is not always true. For example, you can use ulimit to set an upper bound on the amount of address space the process can consume. Also, on a 32-bit system it is easy to imagine a process trying to allocate over 4G of memory and running out of address space.
jes@themisto:~$ cc -o alloc alloc.c
jes@themisto:~$ ulimit -v 100000
jes@themisto:~$ ./alloc
malloc returned NULL after allocated 94227K
EDIT: Ubuntu 15.04, x86_64, Linux 3.19"oh yes it does"
"oh no it doesn't"
demo of malloc failing
"alloca() won't fail!"
Stop moving the goalposts. Besides, alloca() tends to crash your program a lot sooner as stack allocation is usually very limited, especially in a shared-memory multithreading context. Hardly anyone uses alloca().
One reason why overcommit is popular is because of fork(). Imagine a process that does malloc(lots of memory), then forks and execs a tiny program (/bin/true or something like it). If the fork() call succeeds, this means the OS has guaranteed that all the memory in the child process is available to be written over. i.e. it has had to allocate 'lots of memory' x 2 in total, even if only 'lots of memory' x 1 will actually be used.
Without overcommit, fork() can fail even if the system won't ever use anywhere close to the limit of RAM+swap space.
That's a non-sequitur.
If you don't check for NULL return, you have undefined behaviour--that is: your program might end up doing just about anything. If the OOM killer kills you, all visible effects of your program will still be perfectly consistent with the semantics of your program. And you have to be able to safely deal with that scenario anyhow, as much more fundamental resources such as electricity might run out at any time as well.
Imagine your program is processing a folder of emails. If the program is halted part-way through (or if you yank the power cable out) then the mail folder will be left in an inconsistent state.
If the program runs in an OS that does not overcommit memory, it is possible to write code that checks every malloc() and, if it hits a memory limit, you could shut down gracefully, fixing the mail folder so that its state is correct.
If the OS overcommits, then even if you checked every malloc() call, your program might die at any time because of a SIGSEGV or the OOM killer nuking it. There's no way to tidy up an incomplete run.
It's nothing to do with undefined behaviour, or dereferencing NULL pointers.
However, if you don't check for allocation failures, you get undefined behaviour, which you indeed cannot handle safely ... other than by avoiding the undefined behaviour in the first place by checking for allocation failures.
Also, if you check for allocation failures, you won't get a SIGSEGV. SIGSEGV is for invalid virtual addresses, not for lack of resources.
Your point being? That it's untechnically correct?
> The solutions you mention are non-trivial. It is really easy to make a mistake. While it may not be the operating systems fault at the end of the day, designing a system like this is setting up the overall experience for failure. A well designed system should make it easy to get right and not the other way around.
Which is all kindof true, but doesn't change that those APIs are the way they are, for historical reasons. Just because some API is broken, doesn't mean your code will work correctly when you pretend it's not broken.
edit: just to be sure: yes, the special-casing of zero-sized allocations is somewhat of a bug, which we still have to live with. Other than that, the API is actuall perfectly fine--if you don't check for allocation failures and you can't handle it if your program gets interrupted at any point, that's just a bug in your code, there is no useful general way to abstract those problems away.
Allocation failures can and do occur outside of malloc() when overcommit is the memory policy. Malloc might return memory that it believes is valid, but when you try to write to it later on, the OS discovers that it has no free space in the swap and no spare pages to evict. Result: process death (SIGSEGV? SIGBUS? not sure, but it doesn't matter)
Did anyone claim otherwise?
> Whether that is a SIGSEGV or some other signal, or just the heavy boots of the OOM killer nuking your process, doesn't really matter.
It does, because you can catch SIGSEGV.
> The key point is that your program cannot handle them. It is impossible to handle every case.
It can, and it has to, if you don't want it to be defective. (Where "handling" does not mean "continue execution", but "don't corrupt persistent state"--you cannot continue execution with insufficient resources anyway, there is no way around that).
> Result: process death (SIGSEGV? SIGBUS? not sure, but it doesn't matter)
SIGKILL. See above.
Have a system call that allocates memory to a process or process group, but doesn't assign it. This call will fail if overall memory usage is too high.
When the process tries to get memory, it will only tap into that reserved allocation if the system is out of memory.
The OOM killer will never kill processes that have reserved spare memory.
The result? Long-running daemons with stable memory usage can be relatively easily protected against the OOM killer. Ideally the OOM killer never runs amok. But in the real world, it would be nice to be able to limit the damage if it has, and guarantee a usable shell for a superuser while it is going on.
Also, if memory is guaranteed to be available at a later point, it cannot really be used for anything, as there is nowhere to put the contents of that memory the moment the process it's reserved for wants to use it.
But there are two APIs that can be used to implement the same goal: Linux has a procfs setting "oom_score_adj" that can be used to decrease the risk of a particular process being killed (commonly used for sshd for obvious reasons) and there are the mlock()/mlockall() syscalls that you can use to make sure some address space is backed by actual RAM rather than potentially swap.
It is true that the reserved memory could not be used for anything. This is a feature. The memory has been reserved for emergencies, and will keep part of the system usable when everything else goes belly up.
The two alternate methods that you specify are not as useful.
I cannot with oom_score_adj have a shell that can be logged in to and be entirely usable during a fork bomb. Yes, you can ssh in. But good luck running arbitrary commands.
With mlock()/mlockall() I can guarantee that a process is responsive during heavy memory pressure, but good luck if it needs to allocate more memory.
And neither system makes it particularly easy to set things up so that the OOM killer will avoid killing all daemons whose memory usage remained stable during memory pressure. (And that is the biggest problem with the OOM killer, that it killed random things you'd have preferred stayed up.)
But that's due to PID exhaustion, not due to memory exhaustion.
> With mlock()/mlockall() I can guarantee that a process is responsive during heavy memory pressure, but good luck if it needs to allocate more memory.
Well, good luck if the process needs more memory than it had reserved in your scheme. Obviously, you have to mlock() sufficient memory for peak need.
> And neither system makes it particularly easy to set things up so that the OOM killer will avoid killing all daemons whose memory usage remained stable during memory pressure. (And that is the biggest problem with the OOM killer, that it killed random things you'd have preferred stayed up.)
No, it doesn't kill random things, it kills things that are easy to reconstruct (little CPU use so far) and that have lots of memory allocated, and where the oom_score_adj doesn't tell it to spare the process for other reasons.
As for the OOM killer, its logic has changed over time. All that I definitely know is that if a slow leak in application code causes Apache processes to grow too quickly over time, it is usually a good idea to reboot the machine because you never know what random daemon got killed before the actual culprit.
Of course you also need to fix application code...
The logic that is applied probably works well in the desktop/developer case where what went wrong probably went wrong recently and there is a person who can notice. That isn't the context where I've usually encountered it.
That isn't the case anymore, at least on Linux and BSD. fork() uses a copy-on-write scheme so the OS only allocates/copies the parent memory space if the child attempts to write to it.
The point was that after you fork() in a copy-on-write scheme, the OS now has promised that more memory is available to write on than may actually exist. If the OS avoided overallocating, it would have to right there and then reserve lots of memory (without necessarily writing to it) just to be sure that you wouldn't run out at a later date.
the allocators didn't necessary even meet the alignment specification either... even if the documents said they did. i remember reading that malloc on windows would return things on 8-byte boundaries only to find pointers ending in 4 (rather than 0 or 8) coming out of it...
and as many have pointed out returning null may or may not happen in a number of situations. i've seen it happen enough times when i ask for too much memory to believe that it is a useful fail case to check for...
That's what the author's resize function expected, which is why his patch is to use free.
In any case, if I'm compiling with -ansi or -std=c90 on a GCC-based platform that is hosted, the library had better behave in the C90 conforming manner. Only if if -std=c99 is used should it take the liberty with realloc(ptr, 0) being like malloc(0).
No, allocating zero bytes is a perfectly reasonable thing to do, as long as you remember that NULL is not necessarily an allocation failure.
if (((buf = malloc(buflen)) == NULL) && (buflen > 0))
goto OUTOFMEMORY;
is perfectly good code.... but i agree that the code is safer if you check for bad cases.
i can't see why allocating zero ever is more sensible than not doing it at all (and therefore avoiding even having to know about this problem, much less deal with it).
You can have it so that "my_malloc" always returns NULL on a zero length object. Or, if you prefer, so that it always returns a unique pointer:
void *my_malloc(size_t size)
{
return size ? malloc(size) : 0; /* option A */
}
void *my_malloc(size_t size)
{
return malloc(size ? size : 1); /* option B */
}
And you can make my_realloc have the nice freeing behavior: void *my_realloc(void *ptr, size_t size)
{
return (ptr && size)
? realloc(ptr, size)
: !ptr
? my_malloc(size)
: (free(ptr), 0);
}
Note how we still carefully implement a parallel requirement to the one in ISO C. Namely that my_realloc(NULL, size) calls my_malloc(size), regardless of size, just like realloc(NULL, size) is required to be equivalent to malloc(size). We want all the routines in this allocator to be drop-in replacements such that any code which is written to the ISO C allocator spec can be blindly retargetted to use them.Pretty much any serious C program, if indeed it doesn't have an entire allocator of its own, at least wraps the standard one, for the sake of more uniform behaviors, as well as easy retargettability to embedded scenarios. (Not only embedded machines, but say, embedding in a larger application, where you are told "you must use this table of funtion pointers as your allocator" and it doesn't quite look like malloc: zero allocations are not allowed, realloc is missing, ...)
In this particular program, realloc is wrapped under a function called resize; resize had to be patched not to rely on realloc(nonnull, 0), but users of resize don't have to change. The programmer doesn't have to look for fifty uses of resize to fix them. However, it's worth it to wrap the functions at a lower level anyway, with an identical API.
The C library is far from perfect. Do not expect consistency, completeness, beauty, symmetry and so on.
There are worse things in it than realloc(ptr, 0) not behaving in the neat way that you would like. Oh such as:
isalpha(str[42]); // str is char array
being undefined behavior if char happens to be signed, and str[42] is negative, because isalpha takes an int argument which is expected to hold a positive byte value [0, UCHAR_MAX), and not char value. Whoooops! void *my_realloc(void *ptr, size_t size)
{
return size ? realloc(ptr, size) : (free(ptr), 0);
}
According to the C standard, realloc(NULL, size) is guaranteed to be equivalent to malloc(size). From n1548 7.22.3.5 para 3:> If ptr is a null pointer, the realloc function behaves like the malloc function for the specified size.
Also, free(0) is well-defined. From n1548 7.22.3.3 para 2:
> If ptr is a null pointer, no action occurs.
> Unfortunately, some people are not sane, and have decided that it's equally valid for realloc(p, 0) to return NULL meaning-failure and not free p.
What you want is the following semantics:
char *my_old_vector = ... ; /* We created / allocated in some way. */
add_important_stuff_to(my_old_vector);
char *new_vector = realloc(my_old_vector,size+100);
if (new_vector == NULL) {
/* I still have the old vector and can recover. */
}http://blog.httrack.com/blog/2014/04/05/a-story-of-realloc-a...
Looking who wrote the text, I also respect cperciva and believe he must have some good reasons and I'd be glad read which use cases he had, more than just "what's wrong with different reallocs." Because I'm not surprised that the corner cases aren't to everybody's (or even anybody's) satisfaction. It's C. Less is more and all that.
But in which use cases are frequent reallocs actually needed, so much that you can recognize the performance impact? I'd really like to know, as I personally never had such problems. When the single allocations were too expensive I've just used some kind of memory pool. For small stuff realloc is still more expensive than just a few instructions on average when some pool is used.
$x .= $_ while (<>);
I believe this has been fixed now, but perl used to realloc for each append operation, which resulted in O(N^2) time complexity if realloc didn't operate in-place.At worst, using realloc produces the same results as malloc/memcpy/free. At best, it might save a memcpy. No harm in giving it that flexibility.
Obviously though, there is a huge advantage to having a single layer of heap management and letting the heap management algorithm have the best insight in to how memory is being used and needed. Rolling your own realloc on top of the heap manager is as likely to create new inefficiencies as remove them.
> it shouldn't have corners
It's not how C traditionally worked. Almost every function was knowingly made to be not "too good." I see modern programmers expect something else, what even 100 times slower languages not always give. Special cases have different possible treatments, and as soon as the exist nobody can make something that would please everyone.
Ah, I misunderstood what you meant by "doing it in your own function". Yes, wrapping the semantics will work, though doing it right and without imposing some overhead takes a bit of work.
> It's not how C traditionally worked. Almost every function was knowingly made to be not "too good." I see modern programmers expect something else, what even 100 times slower languages not always give. Special cases have different possible treatments, and as soon as the exist nobody can make something that would please everyone.
As the original post pointed out, traditionally it didn't work like this. This is how it has evolved.
I did use poor semantics though, as "undefined behaviour" has a specific meaning in C that is of course how the language was defined. What I meant was "ambiguous or unwanted behaviour". Having a clear behaviour for a particular case, even if it is just returning an error code, is fine by me.
(While it's true that dereferencing a pointer to a zero byte allocation should probably trigger undefined behavior that doesn't mean that it shouldn't be possible to do the != NULL test to see if the allocation was successful.)
edit: turns out that no, neither C11 nor C++11 guarantee that, it was probably historical behaviour.
AFAIK if you try to do this with even a fairly recent Microsoft libc it won't work. I haven't checked the standards and history on this one, but keep in mind it's been only a recent development that MS gives a crap about C99.
The idea that C89 did not mandate this makes sense to me because I remember now-obsolete Unixes not liking this either.
[Edit: The MS documentation claims that it supports realloc(NULL, x) going back quite a while. I know I've been bitten by it not doing that in the current century, however...]
C89 did mandate this:
> If ptr is a null pointer, the realloc function behaves like the malloc function for the specified size.
C89 also mandated that realloc(ptr, 0) be equivalent to free(ptr), which was removed in C99:
> If size is zero and ptr is not a null pointer, the object it points to is freed.
If realloc(p, 0) is a synonym for free(p), what does it mean for it to "also free the allocation p" in the event of failure? Unpacking it, the statement seems to be saying the same thing as, "if it fails, then what it should do is to not fail".
They might not be dumb; they might be forced into it to maintain compatibility with buggy programs and/or configure scripts (e.g. [1]).
" If the size of the space requested is zero, the behavior is implementation- defined: either a null pointer is returned, or the behavior is as if the size were some nonzero value, except that the returned pointer shall not be used to access an object."
To maintain my sanity as a C programmer, I steer away from "implementation specific behavior" as much as possible, thus I prefer checking if buffer size is non-zero before allocating.
However, I do agree it would be better if there was some other kind of error condition, preferably something like the option type from various other languages, or just an easy way to return multiple values like in Go. But C is not that language.
#include <stdio.h>
typedef struct mystruct {
int x;
char *p;
} mystruct;
mystruct func(int n) {
return (mystruct) { n, NULL };
}
int main(void) {
mystruct v = func(1);
printf("%d\n", v.x);
printf("%d\n", func(2).x);
return 0;
}
Works fine with gcc 5.3.1. Multiple destructuring is (maybe) problematic, but if you really want this.There are embedded systems out there (that have C compilers) that have addresses at 0 actually. 8051 C programmers know what I'm talking about.
IIRC, 0 is totally an address when inside the Linux kernel as well. What else do you call the first block of memory on a system?
Just because in userland the Linux Kernel pages memory to a very high-numbered region does not mean that under all circumstances "0" is an invalid address. The entire concept of "null == 0 == invalid" is an innate falsehood, an abstraction brought about by the malloc function (that a lot of other libraries have decided to copy).
But really, what do YOU call the first physical byte of RAM? Most systems I know of... from 8051 all the way to even Linux Kernel... calls it 0.
http://c-faq.com/null/machnon0.html
It might be that the internal representation used by the compiler is different, but from C a NULL pointer will always be equal to 0. See also:
From the standard (http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf), "If an invalid value has been assigned to the pointer, the behavior of the unary * operator is undefined." In a footnote, it goes on to define what invalid values are. "Among the invalid values for dereferencing a pointer by the unary * operator are a null pointer, an address inappropriately aligned for the type of object pointed to, and the address of an object after the end of its lifetime."
And null is indeed defined to be 0. Also from the standard: "An integer constant expression with the value 0, or such an expression cast to type void *, is called a null pointer constant. If a null pointer constant is converted to a pointer type, the resulting pointer, called a null pointer, is guaranteed to compare unequal to a pointer to any object or function."
Indeed. It is undefined within C. It is an abstraction that C builds on top of. Dereferencing 0 on the vast majority of embedded systems (or kernel-level code) is very well defined.
This is just one of the areas where C has a "leaky abstraction".
Frankly it's the dumbest misfeature from C++ that should have had nullptr from day one. What's the point of constructing a rigorous type safety system if you're going to go and break it with magic integers that can act like pointers.