SuperMalloc: A Super Fast Multithreaded Malloc for 64-bit Machines
conf.researchr.org
conf.researchr.org
Your papers were very valuable while doing my undergraduate study. Is Hoard neck and neck with jemalloc and lockless' allocator nowadays?
Paper seems to be there, too, at https://github.com/kuszmaul/SuperMalloc/blob/master/paper/cp...
Does anyone know if there's a cpuid feature flag for HTM? Do all haswell-generation SKUs support the feature?
The author of this paper does not use lock-free techniques, which our allocator used - I'm curious if using lock-free algorithms would have changed the author's design, or improved the performance. I do think that despite the similarities that Streamflow has to TBBmalloc, Streamflow is not susceptible to the same kind of memory blowup. The problem, as described in the recent paper:
"TBBmalloc can have an unbounded footprint. One case was documented by [40]. In this case, one thread allocates a large number of objects, and a second thread then frees them, placing them into the first threads foreign block. If the first thread then does not call free(), then the memory will never be removed from the foreign block to be reused. There appears to be no easy fix to this problem in TBBmalloc, since the thread-local locking policy assumes, deep in its design, that every thread calls free() periodically."
Streamflow avoids this problem by putting the remote-free block check on the malloc path. That is, when allocating memory, you always check if other threads remotely freed memory for you. Basically, if you're continually allocating memory, you're also continually checking to see if you should clean your memory that other threads freed for you. You can see this in action: https://github.com/scotts/streamflow/blob/master/streamflow....
All of the above is just to add some background - I look forward to really digging into this paper over the holidays.
(In case anyone actually follows my comments, they may know I currently do research and development for IBM Streams. That has zero relationship to this memory allocator I worked on early in grad school; it's just an odd coincidence of project names.)
An allocator that is designed for scalability from the ground up. Similar design to Streamflow (based on so-called spans) that eagerly returns memory (latency-aware) and a backend that also makes use of fragmenting virtual memory, which is plentiful available on 64bit systems.
The micro benchmarks are great, the whole program level benchmarks show nothing earth shaking. But I'll take ok performance and half the code any day.
In a 32bit world, taking 1/4 to 1/8 of the address space would be impolite. In a 64 (or 48) bit world it doesn't matter.
SPARC moved to a 64-bit address space a long time ago. 32 terabytes of memory in a single server? Sure, why not.
But as the other poster pointed out, this really isn't an issue given the approach they've taken, even on x86.
The author is giving you the option to omit the MIT text and have a pure GPLv3 if you wish to restrict access to your software in that way.
The MIT license must be included in the provided software, but it isn't a part of the new license.
I think tcmalloc vs. jemalloc is basically a wash. One is written by google people and the other is written by facebook people. Facebook is much more forward about their open source project than is Google, so more people have heard of jemalloc.
And the other side of what I was trying to say is that for Google it's largely an issue of personality. tcmalloc is Google's malloc. JE works for Facebook. If JE worked at Google I'm sure Google would just use jemalloc.