Show HN: Mmm – manual memory management for Go
github.com
github.com
Please elaborate. Do you mean reducing overhead via compressed address space references or offset based?
Nicholas Nethercote has a blog about his work on this on Firefox : https://blog.mozilla.org/nnethercote/ (e.g. https://blog.mozilla.org/nnethercote/2011/11/01/spidermonkey...)
struct foo {
struct bar *bar;
struct baz *baz;
}
into struct foo {
struct bar bar;
struct baz baz;
}
Of course, this is not always possible - but a C program written naively/for encapsulation/flexibly tends to have a lot of places where you can perform such an optimization.The performance win, on modern platforms, comes chiefly from reducing the number of (unpredictable/non-cached) pointer dereferences.
I've been experimenting with a less general segmented & offset addressed approach to deal with the same issues. Quite surprising how much semantics can be encoded in 64 bytes. Performance gains are substantial.
The parenthetical is important. Pointers are just fine as long as you don't dereference them and suffer cache misses as a result.
It's also worthwhile to note that what you described is not always a worthwhile optimization, even if you can do it. If, for example, you traverse an array of the objects containing a sub-object frequently and follow the pointer to that object only rarely, it's often worthwhile to allocate it separately to improve cache locality of that traversal. This is even more important if you're doing a lot of structure copies of the outer structure; avoiding copying more bits saves a lot of time. The latter case comes up surprisingly often in my experience, and I've improved performance of code quite a few times with separate allocations and pointer indirection.
if not then the compiler team should rethink their careers...
It must be that the author only cares about GC pauses, i.e. 'nondeterministic' pauses -- not really perf itself?
Regarding overall performances, I'm not really convinced that the reflect calls affect them that much (there's 2 reflect calls per Read() or Write() call), but I honestly just don't know, I'd have to bench to be sure. Although, if it turns out to actually be an issue, you can still use Pointer() to keep references to your unmanaged heap and you'll be able to work with your data without GC nor reflection overhead.
You'd still want to reach for mmm only after A: verifying that GC time really is your biggest problem via profiling and B: taking more normal steps to minimize overallocation first, because with value types Go has some tools (if not necessarily "a lot" of such tools, but definitely some) for dealing with that. But if you end up backed against the wall, this may be helpful.
You could also arguably add a C: did you really mean to use a language with manual memory in the first place, or perhaps Rust? Or can you factor just the relevant bit out into such a language and interact via some RPC mechanism back to the Go code base? But as the situation becomes arbitrarily complicated there simply ceases to be a silver bullet.
GC pauses can be anywhere between 300ms to 30 seconds or more when it starts becoming an issue.
In this configuration, each incoming request means allocations inside Go's RPC package [1], which in turn means that a GC pass will be triggered if GC_PERCENT [2] has been reached, which in turn means that the GC will have to scan all of those long lived pointers (Go's GC is not generational), which in turn means a huge peak in response time.
This basically leaves me with three possible solutions:
- hack into Go's RPC package to minimize allocations, which is a huge price to pay just to delay the inevitable
- build my caches in a language that offers manual memory management, then query those via RPC from my Go services; but I don't want to add a new language into the mix
- provide a generic solution for manual memory management in Go, which is where we are now
I know many people won't agree with that, and there are definitely good reasons not to; and still, as far as I'm concerned, minimizing the complexity of my software stack means there's one less thing that I'll have to worry about, and at the end of the day, that is really quite the upside.
Again, it's almost certainly premature optimization to start with that design, but if that's where your optimization leads you, it's not that surprising.
FWIW, I wrote something myself that hits a similar problem, but in a completely different dimension: https://github.com/thejerf/gomempool My problem was that I had an otherwise rather placid program (from an allocation perspective) that liked to allocate buffers for messages that were many hundreds of kilobytes to low numbers of megabytes in size. In normal usage, only maybe one or two of these are ever in use at a time, but I use hundreds per second. In my case, each individual GC was actually not that big a deal, but I was triggering them every few seconds. The GC would see a lot of large allocs, and then successfully clean them up, meaning that the next batch of large allocations would be seen as crossing the threshold again. My stats clearly showed that on this system, once I started pooling my []byte I never even filled up the memory pool itself, and my GCs plummeted so far that I wouldn't even particularly care if they took half-a-second apiece anymore, which they don't. Almost everything other than those large message buffers were stack-alloc'ed anyhow.
Of course, I don't have such calls on production systems; and while concurrent collections greatly reduce the time spent in STW phases, latency spikes are still a very real issue, as I explained here [1]. The monitoring on my production systems leaves absolutly no doubt that those latency spikes are due to garbage collections: once the GC was out of the picture, every thing was flat and fast again.
I do know that you need version 1.5 of Go that was released last August to get the low latency GC. If throughput is an issue some folks adjust GOGC to use as much RAM as they can. If none of this helps file an issue report with the Go team. You seem to have a nice well thought out work around but a reproducible gctrace showing hundreds of millisecond of GC latency on heaps with millions of objects will be of interest to the Go team and might really help them.
I hope this helps.
i still hope there is a better solution out there for memory management than using a GC. they are pretty much golden sledgehammers for this problem and enable lots of very bad practice to occur without leaking (this is both a strength and a weakness though...).