This program uses 1 TiB of virtual address space, and runs fine on a machine with no swap (and much less than 1 TiB of physical RAM).
#include <sys/mman.h>
#include <stdlib.h>
int main()
{
size_t const len = 1024 * 1024 * 1024; // 1 GB
for (int i = 0; i < 1024; ++i) {
char *p = (char *)mmap(0, len, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANON,
-1, 0);
if (p == -1) exit(-1);
}
}
The space won't be associated with any physical object (whether RAM or swap) until you actually touch it.I wasn't precise enough. The case where we needed swap not to error out was in a big mapreduce process, consuming maybe 44GB out of a 48GB host. When we didn't have swap on the box, the big process would exec a small command line process and the exec would fail, there wasn't 88GB of address space available during the brief period between the call to fork and the overlay. Adding a swap device fixed our problem even though it was _hardly_ ever used (MB to single GB).
At least that was my analysis from 10 years ago. I never went full science on it, but I did read plenty of documentation on overcommit and virtual address space.
echo 1 > /proc/sys/vm/overcommit_memory
According to documentation, 0 (the default) means usually allow overcommits according to some heuristic (which was probably failing in your case), 1 means always allow, and 2 means don't allow.Edit: It appears from https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin... that things will go wrong if you try to allocate more VM space in one call than the total free RAM+swap available to the system.
When redis is persisting its state to disk, it temporarily needs 2x the memory because it forks a child that does this in the background. Since fork uses copy-on-write, only the difference between child/parent is actually consumed as extra memory. The rest is "de-duplicated" in a sense.
The choice is then to
1) enable swap, possibly killing performance 2) enable overcommit, possibly killing the redis process
It is sometimes preferable for a service to fail hard and fast than get bogged down with swap. I believe this is the idea behind the "no swap" policy, at least on Linux.
Redis is a good example here because any amount of swapping will kill performance since literally all of its allocated memory is consistently accessed for the purposes of doing a dump. And Linux will always try to swap things out, even with vm.swappiness=1.