Talc – A fast and flexible allocator for no_std and WebAssembly
github.com
github.com
In my console, I have something akin to this:
TRACE client_wasm::plugins::allocation: Memory stats counters=Counters { allocation_count: 165454, total_allocation_count: 18756119, allocated_bytes: 34654828, total_allocated_bytes: 3185258585, available_bytes: 82802636, fragment_count: 5026, heap_count: 1, total_heap_count: 1, claimed_bytes: 118423552, total_claimed_bytes: 118423552 }
I haven't carefully benchmarked dlmalloc (Rust's default WASM allocator, https://github.com/alexcrichton/dlmalloc-rs), but it's nothing special (to my knowledge). The swap to Talc is pretty trivial and it's clear that the author is paying attention to its performance.[1] https://github.com/SFBdragon/talc/issues/26 [2] https://github.com/yvt/rlsf
On my random actions benchmarks (this resembles real allocation patterns somewhat better?):
- 1 thread: Talc is faster than Frusa and System, Frusa is comparable to System
- 4 threads: System is fastest, Frusa does about ~half as well, Talc does ~half as well as Frusa
Our benchmarks agree on the Frusa vs Talc comparison.
Benchmarks aside, Frusa seems neat. In particular, I had some misconceptions about how to tackle concurrency in Talc which Frusa's code demonstrates not to be true. I may give writing a concurrent version of Talc another shot soon.
I'm changing up my random-actions benchmark to display results over various allocation sizes, as some allocators do much better than others at different sizes. As a heads up, Frusa takes a large hit at higher allocation sizes. Perhaps tuning bucket sizes or something could help? I'll try to have the benchmarks on GitHub this weekend so you can play around with them, if you'd like to investigate.
This is very common in embedded contexts, where you can take nothing for granted.
1. no_std (like no libc + no malloc in C)
2. no_std + alloc (like no libc + malloc in C)
3. std (like a full libc + malloc in C)
The difference between 1 + 2 is like three lines. The difference between 2 + 3 is a change to the entire standard library. ATM only ESP32 devices support option 3 (they build a standard library implementation on top of FreeRTOS/ESP-IDF).
* file and other I/O, including filesystem ops
* access to system time
* threads
* collections and some other things that require an allocator (not many things actually do in Rust’s stdlib!)
* floating-point functions (the types themselves and builtin operators work fine)
`alloc` gives you `Vec`, `String`, `Box`, `BTreeMap/Set`, ref-counted pointers, and a few `Vec`-derived collections like `VecDeque`. Very annoyingly not `HashSet/Map` though, due to a literally single-line dependence on a system entropy source which happens to not be easily factorable out because reasons.
Currently the `sys` crate implementation is hard-coded into the compiler but eventually you will be able to provide it without modifying the compiler so you can e.g. target a RTOS or whatever.
It looks like that work started really recently actually:
https://github.com/rust-lang/rust/commit/99128b7e45f8b95d962...
It's also the case that such targets often can't take advantage of the advanced features (like multithreading optimisations) that "full fat" allocator provide.
Zig having allocators totally opaque is interesting and you get to learn a lot if you dig deeper. Tons of different designs. Allocators on top of allocators. Tiny ones for small bundles. Non deallocating ones for one shot apps. Stuff you never care about in a GC environment.
I suggest checking it out if it interests you.
I could add these benchmarks. They were there at one point in the past, but it's a disingenuous comparison unless the reader understands the particulars of the workload and the particulars of the tradeoffs each allocator makes. Talc will probably beat these allocators in single-threaded allocation, but will suffer under heavily multithreaded loads and does not currently have the system integrations to release unused blocks of memory back to the system (this can be achieved, to a degree via the OOM handler system, but I haven't yet implemented something like this), nor will it be making syscalls like mmap/sbrk at all.
There is the case where you'd want a faster single-threaded allocation pool within a larger application though, which is a case to be made for using Talc when you have access to the system allocator or mimalloc/jemalloc. Perhaps I'll set up something for that.