Allocators in Rust
smallcultfollowing.com
smallcultfollowing.com
Mozilla has been using jemalloc since the release of Firefox 3 in 2008, primarily because having multiple arenas for allocations of different sizes means that you end up with less fragmentation over time. Fragmentation happens when you deallocate ten 100-byte objects that are scattered around in memory, then need to allocate one thousand-byte object. You can't reuse those 10-byte empty spots in memory, because that 1000-byte allocation has to be contiguous. This means that your total memory usage has to grow. Jemalloc puts small allocations into buckets where they'll be closer together, making it more likely that new allocations will be able to fill the holes left by deallocations, and that freeing small objects will eventually empty a memory page, allowing you to return the page to the OS.
I remember that we were all pretty chuffed at the improvements in Firefox's memory usage; it was quite significant and would surely reduce the general perception that Firefox was a memory hog. Just using jemalloc reduced memory usage by 22% (http://blog.pavlov.net/2008/03/11/firefox-3-memory-usage/ has a really amazing graph comparing FF2 and FF3 to IE7).
Rust uses jemalloc for the same reasons; there's just no technical reason not to use it.
If so, I wonder if this is true for every language, then? For example, for C++, should gcc and clang emit binaries with jemalloc linked in?
1 - https://sourceware.org/ml/libc-alpha/2014-10/msg00419.html
Besides, different applications work best with different allocators so it's definitely a choice the author of a program needs to make.
I'm embarassed to link this twice in a row now, but at one point I played around with the perf of a couple allocators, including glibc's ptmalloc. I wound up answering exactly this question, at least for some simple use cases: http://www.andrew.cmu.edu/user/apodolsk/418/finalreport.html.
Comments in ptmalloc sources say they basically stick to the algorithm described here: http://gee.cs.oswego.edu/dl/html/malloc.html.
Famous last words before someone creates a bug-ridden, slow replacement. See also the papers of Emery Berger et al.
* improved cache locality
* reduced false sharing
* use of multiple arenas for reduced lock contention
[O] http://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemall...
Paper here: http://people.cs.umass.edu/~emery/pubs/berger-asplos2000.pdf
Hoard replacement allocator here: http://www.hoard.org/
(Note that Hoard's design has evolved considerably since 2000; the Mac OS X allocator is based on an older version.)
Wouldn't the distribution of the allocations matter too, switching between the two allocators frequently showing the worst case and reuse of one of two minimising the effect.