Unless I misunderstood that the default Rust allocator, with high request bodies and concurrency, is always going to suffer unfixable heap fragmentation like displayed in the article?
Unless I misunderstood that the default Rust allocator, with high request bodies and concurrency, is always going to suffer unfixable heap fragmentation like displayed in the article?
* Switching from libc malloc to jemalloc
* Switching from libc malloc to tcmalloc (dating myself a little bit)
* Switching from libc malloc to mimalloc
* Switching from jemalloc to mimalloc
* Switching from jemalloc to libc malloc
* Switching from mimalloc to jemalloc
Possibly others; I only want to list cases I'm 100% certain of.Heap fragmentation is just a reality of some allocation patterns without a GC runtime.
One certainly can (and, in some cases, should) make their application more allocator-friendly, but - aside from some often-low-hanging fruit - this is a time-intensive process involving a bit of, for lack of a better word, arcane knowledge (I should inline all my fields and allocate on the stack as much as possible, right? Yes, well, except ...)
If you already have a halfway decent benchmark suite or workload generator, which you'll want for other purposes anyway, it's often a lot quicker to just try a few other allocators and select the one that handles your workload best.
If you think of tcmalloc as an old crusty allocator, you've probably only seen the gperftools version of it.
This is the version Google now uses internally: https://github.com/google/tcmalloc
It's worth a fresh look. In particular, it supports per-CPU caches as an alternative to per-thread caches. Those are fantastic if you have a lot more threads than CPUs. I haven't checked if it's been adapted for the latest upstream kernel API, but there's also the idea of "vcpu"-based caches: basically rather than a physical cpu id, it's an (optionally per-numa-node-based) dense id assigned to active threads, so that it still works well if you have a small cpu allocation for this process on a many-core machine.
Jemalloc may more recognition in the broader community, but the largest workloads seem to be running mimalloc / tcmalloc (I don’t know what Facebook uses internally). The libc malloc probably has even more users than either as it’s the default allocator for iOS and Android.
tcmalloc at least gets internal changes regularly synced to github (since just a couple years ago iirc), vs. the ancient gperftools snapshot that's more widely known. And there are many other projects of potential interest to the outside world (Fibers...) that have been mentioned publicly but not open sourced at all.
It is also a reality with a GC runtime, even a compacting one. Tracing GCs have acceptable performance only if the available memory is many times larger than the memory used. Technically this maybe isn't called fragmentation, but the effects are just as bad - the application uses much more memory than really needed.
It's possible that the glibc allocator contains some simple bug. It's more likely that it doesn't contain any simple bugs, but makes different tradeoffs to jemalloc, which make it less suitable to this particular slice of applications.
There might not be anything wrong with the default allocator, it just isn't the best suited for that particular use.