Programs compiled by Go 1.11 allocate an unreasonable amount of virtual memory
github.com
github.com
The nice thing about virtual memory is that it's, well, virtual. It costs you almost nothing until you've touched it. (Fun exercise for the reader: measure the kernel overhead for an unused 1 TiB VMA.) But creating huge spaces--that terabyte mmap wasn't theoretical--that stay untouched is hugely algorithmically useful, especially for things like malloc implementations.
Why does it bother people? Two reasons. First is mlock to avoid swap. This is solvable in much better ways--I'm a fan of disabling swap in many cases anyway. Second is that, absent cgroups, it's difficult to put hard limits on memory usage in Linux. So people, looking under the streetlight, put limits on virtual usage, even though that's not what they care about limiting! Then they get angry when you break it. My refrain here, as in many cases (see for example measuring process CPU time spent in kernel mode): "X is impossible" doesn't justify Y unless Y correctly solves the problem X does.
(I spent years in charge of a major memory allocator so this is a battle I've fought too many times.)
package main
func main() { for { } }
Using 100mb+ of memory?
It's not using them.
On modern OSX, processes get something like 2.4GB vmem by default, even if they do nothing.
#include <unistd.h>
int main() {
for(;;) {
sleep(10);
}
}
is reported as 2377M VMEM by top/htop, on 10.11.You're right that it doesn't cost anything, other than the risk that a process can cripple your machine using its overcommitted memory mapping. And so the kernel has protections against this, which should deter language runtime developers from doing this.
And let's not forget that MADV_DONTNEED is both incorrectly expensive on Linux and ridiculously expensive compared to freeing memory and reallocating it when you need it. Bryan Cantrill ranted about this for a solid half an hour in a podcast a year or two ago.
Also, I assume the crippling you’re talking about here is just the ability to rapidly apply memory pressure? Otherwise I’m very confused.
Sorry, I didn't phrase it well. MADV_DONTNEED is significantly more expensive than most ways that memory allocators would "free" memory. This includes just zeroing it out in userspace when necessary (so no need for a TLB modification), or simply unmapping it and remapping it when needed.
> Also, I assume the crippling you’re talking about here is just the ability to rapidly apply memory pressure?
Right, and if the memory is overcommitted then you can cause OOM very trivially because you already have more mapped pages than there is physical memory -- writing a byte in each page will cause intense memory pressure. Now, this doesn't mean that it would kernel panic the machine, it just means it would cause issues (OOM would figure out what process is the culprit fairly easily).
This is why vm.overcommit_ratio exists (which is what I was talking about when it comes to killing a machine) -- though I just figured out that not all Linux machines ship with vm.overcommit_memory=2 (which I'm pretty sure is what SUSE and maybe some other distros ship because this is definitely an issue we've had for several years...).
There's also RLIMIT_AS, which applied regardless of overcommit_memory.
As an aside, it seems like an apples and oranges comparison to compare “freeing” by zeroing (which doesn’t free at all) to MADV_DONTNEED. I’m also pretty sure that munmap will be much slower than MADV_DONTNEED, or at least way less scalable, given that it needs to acquire a write lock on mmap_sem, which tends to be a bottleneck. It does seem like there’s a lot of opportunity for a better interface than MADV_DONTNEED though (e.g. something asynchronous, so you can batch the TLB flush and avoid the synchronous kernel transition).
Once the cgroup OOM bugs get fixed, amirite? :P
> It does seem like there’s a lot of opportunity for a better interface than MADV_DONTNEED though (e.g. something asynchronous, so you can batch the TLB flush and avoid the synchronous kernel transition).
The original MADV_DONTNEED interface, as implemented on Solaris and FreeBSD and basically every other Unix-like does exactly this -- it tells the operating system that it is free to free it whenever it likes. Linux is the only modern operating system that does the "FREE THIS RIGHT NOW" interface (and it's arguably a bug or a misunderstanding of the semantics -- or it was copied from some really fruity Unix flavour).
In fact, when jemalloc was ported to Solaris it would crash because MADV_DONTNEED was incorrectly implemented on Linux (and jemalloc assumed that MADV_DONTNEED would always zero out the pages -- which is not the case outside Linux).
> As an aside, it seems like an apples and oranges comparison to compare “freeing” by zeroing (which doesn’t free at all) to MADV_DONTNEED. [...] I’m also pretty sure that munmap will be much slower than MADV_DONTNEED.
This is fair, I was sort of alluding to writing a memory allocator where you would prefer to have a memory pool rather than constantly doing MADV_DONTNEED (which is sort of what Go does -- or at least used to do). If you're using a memory pool, then zeroing out the memory on-"allocation" in userspace is probably quite a bit cheaper than MADV_DONTNEED.
But you're right that it's not really an apt comparison -- I was pointing out that there are better memory management setup than just spamming MADV_DONTNEED.
http://www.bsdnow.tv/episodes/2015_08_19-ubuntu_slaughters_k...
http://www.bsdnow.tv/episodes/2015_11_23-the_cantrill_strike...
(I think the second one)
That you have to keep fighting this battle is an indication that people's needs (or desires) aren't being well met.
That said, I don't find it unreasonable at all. Just reserving some bits in the address space isn't unreasonable. It makes the real allocation code simpler.
> The significant increase in virtual memory usage is usually not an issue, however security sensitive programs often lock their memory, causing a far greater performance degradation on low-spec computer hosts.
what does locking memory mean in this context? For what purpose do security sensitive programs lock their memory? what is the performance degradation that happens with low-spec computers?
You should always lock memory if you're going to be storing crypto keys, etc. since once the pages are swapped to disk you're vulnerable to someone pulling the swap partition out and reading it.
1. You don't have to lock all your memory (although it may be hard to capture all the intermediate buffers if you don't).
2. You still need to clear the memory buffers after they're no longer going to be used, otherwise other processes can read /proc/kcore, etc (or cool your RAM and extract it and put it another system)
3. It is possible to encrypt the swap partition with a randomly generated key at boot
I think the person is referring to the mlock() and mlockall() functions (or equivalents on other OS), which keep pages resident / prevents pages from being paged out. Forces them to remain in RAM.
It can be used to, e.g., prevent an encryption key or password from being swapped out to disk, where it might then be recoverable. (Personally, this is why I encrypt swap.)
> what is the performance degradation that happens with low-spec computers?
Locking a larger portion of RAM means less room for the OS to page out unused pages and free up the space for other programs.
While one can try to selectively lock buffers with sensitive data with mlock(), you have to be sure they aren't copied into other buffers that aren't locked (and could thus be subsequently paged out). If you're writing a UI program that displays or receives those in a widget, this might be harder (you might not have access to the internal buffer of the widget, as it is an "implementation detail" of your library), and locking the entire process might be a simpler solution (albeit being a bigger hammer).
But, in the context of this overall discussion, I think it's important to keep in mind that when a process allocates a large amount of virtual memory, it does not automatically allocate any physical memory. So a process allocating a lot of virtual memory up front should not impact other processes which have locked some of their memory into physical memory.
It will, if you mlock() it, I believe. The manual page notes "real-time processes" as a main user of mlock() (the other being the cryptographic uses I hinted at); it cites their use case as locking the page to avoid delays due to paging during critical sections. In order for that to work, the OS would need to bring the pages in, at the time of locking; so at that point, a large virtual allocation becomes equivalent to a physical one.
What it does is keep a canonical version of the page in memory. That's useful being able to deterministically touch a piece of memory, but it doesn't really help you as far as making sure the page never touches disk.
> Memory locking has two main applications: real-time algorithms and high-security data processing. Real-time applications require deterministic timing, and, like scheduling, paging is one major cause of unexpected program execution delays. Real-time applications will usually also switch to a real-time scheduler with sched_setscheduler(2). Cryptographic security software often handles critical bytes like passwords or secret keys as data structures. As a result of paging, these secrets could be transferred onto a persistent swap store medium, where they might be accessible to the enemy long after the security software has erased the secrets in RAM and terminated. (But be aware that the suspend mode on laptops and some desktop computers will save a copy of the system's RAM to disk, regardless of memory locks.)
I would be rather disappointed if a hypervisor swapped out my guest (at least, in a context like AWS; I suppose if you're just running qemu on your laptop, that's a different matter), but I hadn't considered that either, and it is certainly possible.
> All pages that contain a part of the specified address range are guaranteed to be resident in RAM when the call returns successfully; the pages are guaranteed to stay in RAM until later unlocked.
So it doesn't matter if the memory is dirty, as long as it's marked as locked.
The current Linux man page gives a bit more insight:
mlockall() and munlockall()
mlockall() locks all pages
mapped into the address space
of the calling process. This
includes the pages of the code,
data and stack segment, as well
as shared libraries, user space
kernel data, shared memory, and
memory-mapped files. All
mapped pages are guaranteed to
be resident in RAM when the
call returns successfully; the
pages are guaranteed to stay in
RAM until later unlocked.
The flags argument is con‐
structed as the bitwise OR of
one or more of the following
constants:
MCL_CURRENT Lock all pages
which are currently
mapped into the
address space of
the process.
MCL_FUTURE Lock all pages
which will become
mapped into the
address space of
the process in the
future. These
could be, for
instance, new pages
required by a grow‐
ing heap and stack
as well as new mem‐
ory-mapped files or
shared memory
regions. $ cat wheres-the-ram.c
#! /home/rkeene/bin/c
#include <sys/mman.h>
#include <stdlib.h>
#include <stdio.h>
int main(int argc, char **argv) {
unsigned char *buffer;
int mla_ret;
mla_ret = mlockall(MCL_CURRENT | MCL_FUTURE | MCL_ONFAULT);
if (mla_ret != 0) {
perror("mlockall");
return(1);
}
buffer = malloc(1024LLU * 1024LLU * 1024LLU * 32LLU);
if (!buffer) {
perror("malloc");
return(1);
}
buffer[0] = 1;
buffer[1] = buffer[0];
buffer[2] = buffer[3];
puts("Success !");
return(0);
}
$ sudo ./wheres-the-ram.c
Success !
$ free -g
total used free shared buff/cache available
Mem: 14 1 5 0 7 13
Swap: 0 0 0
$Operating systems usually cap the amount of locked pages to prevent the system from being DoS'd; on Linux it can be quite low (16kb).
In all the code I've written using locked memory, these limitations have forced me to use separate arenas for the locked/sensitive memory because of its scarcity. In general, it would be incompatible with Go's garbage collected heap unless the GC heap has the concept of "sensitive" objects and pools them in locked memory (which, AFAICT, it doesn't); or the heap limited itself to the locked memory limit (impractical)
If you allocate and maintain unmanaged locked memory yourself in Go it shouldn't matter if the 1.11 runtime uses more virtual memory since you've separated yourself from the problem by going your own route.
The ideal for me would be a function that marked some memory as "these addresses are taken, do not give them out to malloc, or anything else", but which still required me to actually "ask" for the memory before using it, so it didn't look like I was using 64gb of memory at startup. Is that possible?
Reserves a range of the process's virtual address space without allocating any actual physical storage in memory or in the paging file on disk.
https://msdn.microsoft.com/en-us/library/Aa366887(v=VS.85).a...
Perhaps mmap() can achieve the same on Unices?
Reading that, I think overcommitment is to determine the kernel's behavior when you try to allocate more virtual memory than physical memory that is present on the system. That is a different (but related) concern from the fact that mmap() will allocate virtual memory but not physical memory. (That is, mmap() reserves locations in the address space, but you don't have any physical memory backing it until you use that memory.)
Personally, my feeling is: why does it matter if it "looks like" a process is using a lot of memory? That is, why does it matter if a process allocates a lot of virtual memory up front? It's not consuming physical memory, it's just updating some bookkeeping in the kernel. I know people feel uneasy about seeing large values for virtual memory, but... they shouldn't.
So you want to reserve memory, basically the way malloc does, depriving any other process of that memory, but you then want to reserve it again, and also somehow hide that the memory has been reserved?
What exactly is the benefit except making it not look like you're using the memory? Actually, what's the benefit of that in itself?
(I don't see any request for different physical memory allocation, so I assume that would still be handled by page faulting in the kernel.)
People do it all the time in real life, with area codes, zip codes, case numbers, etc.
For example, a long street full of strip malls in California will often have street numbers 50 apart to allow them to remain sequential even after new developments.
No one would complain that this is a wasteful use of precious street numbers, or that it deprives other streets of those numbers.
Imagine the nightmare if you periodically had to renumber all the buildings instead, like computers routinely have to.
One can already do this with realloc() (https://linux.die.net/man/3/realloc). And if you don't want to use the malloc family, but instead want to manage your virtual memory yourself, you can easily allocate large sections of virtual memory and then manage it yourself - which would include enabling behavior such as growing arrays. (Even though you're really just re-implementing realloc-like behavior yourself.)
And to make sure we're on the same page (ha), I want to reiterate that there is a big difference between virtual and physical memory. You can allocate virtual memory without allocating physical memory.
The original point was entirely to avoid scaring newbies who can't tell virtual from physical, while still being liberal with your virtual usage.
Haskell with GHC has a great solution: just allocate 1TB up front. A newbie does not need improved tooling, a reworked memory system or any education to realize that this can't possibly be RAM.
This is surprisingly hard to do in user space. You can tell mmap where to put allocations, but not where not to put them. Also it's hard to control what your Libc's malloc will do.
Sure, it will show up in the "virtual memory" usage of your process. That's just how the virtual memory accounting works.
Nobody that I know of works that way besides some L4 kernels and they only use it as an IPC optimization strategy.
(and unfortunately, at least according to the L4 people, almost half the cost of a context switch - fixing the TLB)
From the documentation:
> Instead of a mapping, create a guard of the specified size. Guards allow a process to create reservations in its address space, which can later be replaced by actual mappings. mmap will not create mappings in the address range of a guard unless the request specifies MAP_FIXED. Guards can be destroyed with munmap(2). Any memory access by a thread to the guarded range results in the delivery of a SIGSEGV signal to that thread.
So it seems like a good fit for your use case, but I've never used this.