Nuke: A memory arena implementation for Go
github.com
github.com
arena := NewSlabArena(8182, 1) // 8KB
var b *byte = New[byte](arena)
var i *int = New[int](arena)
fmt.Printf("Pointer address: %p\n", b)
fmt.Printf("Pointer address: %p\n", i)
Result: Pointer address: 0x14000198000
Pointer address: 0x14000198001
I'm not a Go language lawyer, but I assume this is just immediate UB. OP, it's fine to publish a library without experience in manual memory management, but maybe put a disclaimer in the README?Maybe? But that seems strange, since they seem to have intended it to be used for http servers making per-request allocations.
Come to think of it, the arena is probably still unsound with just ints, because the underlying allocation is just for a `[]byte`, which I don't think is guaranteed to be aligned to 8 bytes. Might be on most platforms, though.
Nginx does this. As does my Passenger application server.
Apache... kind of does this, but it fakes it. It allocates every object individually, and the "arena" is only used for linking all those allocations together so that Apache can free all of those allocations (individually). facepalm
In manual languages like C or C++, you can use these to allocate a fixed set of memory on program init (keeps system resources under control), to get contiguous allocations (friendly for caches), to keep yourself sane for memory (clear start and end to the lifecycle of an object), and to be very performant (if used correctly).
Language selection is really important and I think too many engineers approach it rather willy-nilly and with way too much bias towards what they like or may already know. Both of those are legitimate considerations! But they shouldn't be determinative. You need to calmly and rationally look at all the tradeoffs the languages offer. I think the vast majority of projects that have the sort of performance requirements that require arenas to function could have had that requirement determined from the beginning and the conclusion reached that Go was not a good choice, despite matching on some criteria. If this degree of memory performance is a critical requirement for your project, you're looking at a list of possible languages I could count with one hand, and Go's not on it.
(Though based on what I see in the world right now, the more common problem is people getting a project and grotesquely overestimating the performance they need, like, the guy tasked with writing a web site that will perform up to 5 entire CRUD updates per second at maximum load posting questions about whether they need the web framework that does six million requests per second or the one that does ten million. But both over and under estimating requirements is a problem in the real world.)
I would think not twice, but more like a dozen times about using a package like this. I would need to be backed into it by sheer desperation, some large code base that I simply can not fix any other way, or extract this into its external service/microservice/library for the task, or literally almost anything else, and using it would represent my program reaching the end of its "design budget", if not exceeding it.
And unless the decision was just so far back in the mists of history that it is completely irrelevant now (e.g., the decisions were all made by people no longer on the project), there'd be a postmortem on how we made the mistake of picking an inappropriate language.
I do alot of profiling and performance optimization. Especially with go. Allocation and gc is often a bottleneck.
Usually when that happens you look for ways to avoid allocation such as reusing an object or pre allocating one large chunk or slice to amortize the cost of smaller objects. The easiest way to reuse a resource is something like creating one instance at the start of a loop and clearing it out between iterations.
The compiler and gc can be smart enough to do alot of this work on their own, but they don't always see (a common example is that when you pass a byte slice to a io.Reader go has to heap allocate it because it doesn't know if a reader implementation will store a reference and if it's stack allocated that's bad.
If you can't have a clean "process requests in a loop and reuse this value" lifecycle, it's common to use explicit pools.
I've never really had to do more than that. But one observation people make is that alot of allocations are request scoped and it's easier to bulk clean them up. Except that requires nothing else stores pointers to them and go doesn't help you enforce that.
Also this implementation in particular might not actually work because there are alignment restrictions on values.
https://github.com/Enichan/Arenas
For C#.
Been trying to get into more managed memory in C#, so this might be something good for that.
Arguably, there's less need for arenas in C# in most scenarios than in Go because of easy object and array pooling out of box with ArrayPool<T>, ObjectPool<T> (Sdk.Web workload) and stackalloc/InlineArray and co. You can also just use malloc/free directly with NativeMemory.Alloc/Free instead.
If I am not misunderstanding what this project tries to do, then they are essentially adding another layer, they get large chunks from the allocator and then use those to satisfy allocations. This seems essentially like having a second allocator on top of the existing one. If the existing one does not work well in certain scenarios, there might be some performance to be gained by using a different allocator, even on top of the existing one. But I wonder if this is the best way to solve the issue, this seems more like a workaround than a fix. Would it not be better to tune the existing allocator or make it configurable or even swappable? This of course requires more fundamental changes - language, runtime, compiler - instead of just being a library. Or maybe Go already has facilities to customize memory management?
EDIT: I think I got this wrong, the actual goal seems to be able to allocate several objects in a continuous chunk of memory, i.e. an array of objects - not to be confused with an array of pointers to objects - for improved locality.
They also allow you to have multiple arenas.
A use case is a web server, where each request creates an arena, allocates scratch memory in it, and destroys the arena after handling the request.
It is not. Arena allocators have a fundamentally different API, because they don't allow you to de-allocate individual objects in the arena - everything must be de-allocated at once. For specific workloads, like repeatedly allocating a large number of short-lived objects, this can be a huge speedup, and also significantly reduce memory usage.
https://www.cs.cornell.edu/courses/cs6120/2019fa/blog/immix/