When the memory allocator works against you
glandium.org
glandium.org
>“Import manifest” phase (which, in fact, allocates and frees increasingly large buffers for each manifest, as manifests grow in size in the mozilla-central history)."
I am not familiar with the reference of a "manifest" or importing them at least not in the context of memory allocation. Can someone shed some light on this?
Also right after the author states?
>"To put the theory at work, I patched the python interpreter again, making it use malloc() instead of mmap() for its arenas."
If I understand correctly the author is trying to remove the possibility of entanglement but doesn't the malloc() glibc library routine translate into an mmap() system call? I am not sure how that would help disambiguate memory allocators then.
+of git operations.
It is git, it is made to grow and get into billions of string comparisons. Anyways...
C/c++ would have probably been a better language for his problem.
Simply prealocate as much as needed as a working buffer and eliminate the need to profile libc and patch your interpreter altogether.
Not necessarily, it depends on your allocation patterns. The same can be said for using malloc.
> Not to mention, it has no idea how your program uses memory, such that it is bound to use a generic approach
Ditto for malloc.
> with manual management, you can get real and high performance.
Using a GC does not mean you can't manually manage allocations. E.g. say you are allocating and deallocating a lot of Point objects. You keep a free list of Point objects that you can recycle. If you need a Point, you first see if there are any in the free list. If so, you pop the first one off the list and reuse it. Otherwise you allocate a new one. When you're done with a Point, you push it on the free list.
> Ditto for malloc.
malloc != manual management
> Using a GC does not mean you can't manually manage allocations. E.g. say you are allocating and deallocating a lot of Point objects...
But that's not using a GC. It's manual memory management.
> If you use malloc willy-nilly throughout core algorithms, your runtimes become unpredictable.
First, where did I say "use malloc willy-nilly"? Second, do you have experience with that? I haven't much, but I haven't heard your claim before.
Of course "willy-nilly" should not be done with real-time constraints. But I have written a substantial algorithmic program (no RT constraints) in that style when I didn't know better, and it was very well performing. (probably glibc's malloc was optimized for my allocation patterns).
The main problem I see with that style is that each malloc has memory overhead, and that it leads to unmaintainable code.
Well, that was true when it was first released, but recent versions have gained "native" support to talk to mercurial servers (where a helper program written in C does the talking[1]), although the mercurial code is still used to read bundle2[2] data that comes from the mercurial server, even in that case (that is, only bundle1 is supported without using mercurial code).
I'm moving more and more things to the helper with the ultimate goal of not having any python code left.
1. https://glandium.org/blog/?p=3648 2. https://www.mercurial-scm.org/wiki/BundleFormat2
Also, to clarify, the amount of memory allocated is actually decreasing (in-use bytes), while the memory allocator requests more and more memory from the system (system bytes), presumably because of fragmentation.
Excuse my ignorance, but couldn't it still be a problem of the python runtime? I.e. they use malloc wrong, etc?
Just the other day was reading N. Wirth objecting on using a resource lavishly just because it was cheap. And just now, in the same front page of HN, some guy wants programmers to stop calling themselves engineers because they aren't.
Well, let's build a bridge using all the iron we can get our hands on (iron is very cheap you know); or even better, let's use all the RAM in the computer with this program of ours.
Memory allocation "doesn't fail" so they kill a random program?
Can someone link me to an explanation of this?
That's not quite it. But it does deserve a double-take.
When you request memory, it generally works out fine. But it doesn't actually give you a lock on the pages you need until you request those pages. This leads to weird situations where memory is oversubscribed. When this happens the OS casts one of your processes into the rift. Google for "oom killer".
To me this sounds insane - a basic breach in the contract of malloc. However, it's possible that this is one of those situations that involves a complex tradeoff that you won't grok until you've been the guy building it. It would be interesting to know what other OSs do.
Something that is annoying: as far as I know there is no signal that fires when this happens. Instead, you scrape logs to learn about it. Whereupon you will reboot the box, because you no longer trust its state. You can get a callback when this happens if you use systemtap.
In practice it's less of a problem than you'd think. If you care about the thing you're building, you will have made it host-independent. From the perspective of production software, there's no difference between a rifted process and a motherboard death. If you're using sockets as your IPC mechanism, it doesn't matter why the process died, just that it did.
One is to commit memory on request. If all of your RAM and swap is committed, then malloc will fail. This is probably how we expect it to work, but it can be really wasteful. Some programs will malloc a huge area and only use a small portion of it, and in this scheme, that memory is "in use." It's also worth noting that most programs don't tolerate malloc failure well at all, and will just crash anyway.
Another way is to overcommit, where malloc doesn't actually reserve memory initially but only on first use, but only pause processes when memory fills up. If a process needs more memory and there isn't any more available, that process blocks until more memory becomes available. This means nobody gets killed, but it's not uncommon for the system to just deadlock like this forever once memory fills up.
And the third way is as you describe, to overcommit and kill processes when you run out of memory.
Note that there's no right answer, just different tradeoffs. Linux lets you configure this to an extent as well, so if you don't like the default, you can change it!
I'm curious how committed memory interacts with COW fork semantics. After all, a child can just overwrite the COW/copy of the memory it inherited from its parent without allocating anything else.
But at least then it's the program that doesn't handle failure that dies, not one that does!
...and it would be absurd for bad programs to eat memory and crash? We must have other programs pay for its sins?
Apps are supposed to be built to tolerate being killed while in the background and restoring their state when relaunched, so in theory this will mostly be invisible to the user other than apps taking longer to reload. In practice, lots of apps don't get this right.
The downside is memory allocations that would oversubscribe memory will immediately fail. That means processes won't be able to allocate tens of gigabytes of virtual memory anymore.
If I understand the limitation correctly, this also means that all process's allocated virtual memory can no longer exceed the amount of physical+swap memory in the machine. I think this probably interacts badly with COW memory semantics for child processes: even though children and parent reuse address space, the kernel can't know whether the child will write to the memory it inherited from its parent's memory map. I think if the kernel is being too careful, all "shared" memory would have to be accounted separately, which would severely limit the number of processes you could have open.
(Is that really what happens? That seems awfully limiting. I hope I'm wrong.)
You can almost think of file-backed mmaps as swap, in fact.
Is it just that software is (badly) coded to request lots of memory it doesn't intend to use? The only sensible usecase I can think of for that is oversubscribing VMs, but even there... Doesn't this just crash a VM when there's contention for the memory that's oversubscribed?
I guess that would be okay for the VM case, but why would I want this in a normal usecase? It seems like it's just encouraging/supporting bad memory management in software.
Have I just been lucky that this has never killed a long-running job of mine when I use 90%+ of RAM to do a big job? Or shit, maybe it has and the specific error just was obscured. Ugh, I hate when the system lies by default.
How does that help me as a user?
It's a SIGKILL, which isn't catchable/handleable by the process (how could it be?).
The machine should still be stable after an OOM, just minus a few process. (In my production setups, usually a supervising daemon can restart the now-dead child, and begin the memory leaking cycle anew.)
For example, caching. It's essentially impossible to do caching at the application layer right unless your cache is also persistent with no durability or consistency needs whatsoever and can be mmaped (and you don't mind total I/O trashage). The persistence part also implies that you must be able to store the entire cache somewhere.
What an application would want here is a mechanism to allocate cache memory that the kernel can throw away, if need be, and notify the application of that, so that the app can adjust. That's not impossible to do, it's just tricky - and beyond POSIX.
The difference is that engineers have codified this and call it "safety factor." Your house is built to carry perhaps ten times the load it truly needs to, because it's cheaper to throw in extra materials than it is to build it to hold precisely what it needs. An airplane is built with much thinner margins, because saving materials is far more valuable, and it's well worth the extra time and attention to.
But just because you have enough of a cheap resource (RAM) it is not a given that using it in excess is either optimal or robust.
As it is now, CPU, RAM and storage usage are perfect externalities for developers. The result of their abuse is user frustration, and reduction of the number of things people can do on their computers. That, and wasting user time. Unfortunately, there's no way that would make the developers suffer costs for wasting user's resources, so the only thing that lets us have lean software is engineering competence and human decency (both of which seem to be in short supply).