I also found a simple solution: call malloc_trim() after a GC. This reduces memory usage by 70%.
https://www.joyfulbikeshedding.com/blog/2019-03-14-what-caus...
I also found a simple solution: call malloc_trim() after a GC. This reduces memory usage by 70%.
https://www.joyfulbikeshedding.com/blog/2019-03-14-what-caus...
This means there's a wave of new allocations moving through your address space, and it's leaving behind a fragmented mess. Calling malloc_trim() won't help with the address space fragmentation, it will only free memory pages caught up in the mess. At some point the allocation wave will hit the top of the address space and allocations will start to fail. Usually this is not a problem in 64-bit processes of course, because it will take a very long time to run out of 64-bits, but on 32-bit processes this was a real problem.
This is what the MMAP_THRESHOLD tunable solves. It makes that allocations larger than that many Bytes are served via their own mmap that can be munmapped in independence.
I use env MALLOC_MMAP_THRESHOLD_=65536 to reduce the memory-fragmentation wasted RAM of my program from 6.5 GB to 0.8 GB.
The benefit of this is that you don't have to decide at which points to call malloc_trim(). But it's expected to be a bit slower because mmap() takes a while. Choosing between malloc_trim() vs MALLOC_MMAP_THRESHOLD_ is dual to choosing between GC vs reference counting -- higher memory use for a while and having to choose when to clean up vs higher per-operation cost.
If large allocations are page-aligned (hopefully at 16+ pages), they can be individually unmapped and remapped, and I see no reason why or how individual mmaps could in general result in less fragmentation, other than said silly malloc.
E.g. https://github.com/thestinger/allocator/tree/f42a6c2dffb63d5... (found via Google) explains:
> The Linux kernel also lacks an ordering by size, so it has to use an ugly heuristic for allocation rather than best-fit. It allocates below the lowest mapping so far if there and room and then falls back to an O(n) scan. This leaves behind gaps when anything but the lowest mapping is freed, increasing the rate of TLB misses.
So the libc should be able to do a better job than the kernel here.
And yes, I think it should probably work via MADV_DONTNEED to give mapped pages back to the kernel (and perhaps PROT_NONE to also reduce commit charge when overcommit is disabled, see https://github.com/thestinger/allocator/issues/18).
I'm using `MALLOC_MMAP_THRESHOLD_` because it has a positive effect, but as written on https://news.ycombinator.com/item?id=24244271, I'm not sure why automatic trimming (which should do all of the discussed above) does not work.
According to the docs [1,2] it should be called automatically when the free space exceeds the default M_TRIM_THRESHOLD of 128 KiB.
Is it because of this bug [3] or for another reason?
[1] https://man7.org/linux/man-pages/man3/malloc_trim.3.html
[2] https://man7.org/linux/man-pages/man3/mallopt.3.html
[3] https://sourceware.org/bugzilla/show_bug.cgi?id=14827CPython is less affected than CRuby because CPython has a specialized allocator called obmalloc for small objects up to 512 bytes.
CRuby < 2.6 doesn't have an allocator like this and hits malloc for anything bigger than 24 bytes. 2.6+ can allocate using the "transient heap" which helps but isn't as effective as CPython's obmalloc.