Memory by the Slab: The Tale of Bonwick's Slab Allocator [video]
paperswelove.org
paperswelove.org
In terms of follow-on work, Ryan mentioned the later libumem work[1], but it's also worth mentioning Robert Mustacchi's work in 2012 on per-thread caching in libumem[2]. And speaking for myself, I am indebted to the slab allocator for my work on postmortem object type identification[3] and on postmortem memory leak detection; both techniques very much relied upon the implementation of the slab allocator for their efficacy.
Thanks to Ryan for bringing broader attention to a terrific systems paper -- and truly one that I personally love!
[1] https://www.usenix.org/legacy/event/usenix01/full_papers/bon...
[2] http://dtrace.org/blogs/rm/2012/07/16/per-thread-caching-in-...
I designed a custom allocator before new/delete operator overload for C++ OO app. You can think of the app like MS Word, when you open/create a new doc, one need a lot of malloc(). In my case it usually between a few millions to a few tens/hundreds billion records.
There were a lot of overhead for standard new/delete. After profiling, I end up writing my own allocator with the following property:
* It malloc 1,2,4,8,16,32,64MB at a time. (progressively increase to optimize the app RAM footprint for small and large doc use case.
* All the large block alloc()s are associated with the "Doc/DB".
* When the Doc close, the only freeing a few large block are needed. This change make the doc/db close operations go from 30+ seconds for large Doc/DB to less than 1 seconds.
* I later modified the allocate to get the large block memory directly from a mmap() call. All the memory return are automatically persistent. The save operation also went from 30+seconds for large multiple GB DB to < 1 seconds. (Just close the file and the OS handle all the flushing, etc.)
Without ability to customize memory allocator + pointer manipulation, I can't figure how to get similar performance for similar type of large scale app with Golang, Java, etc.