(* ) http://cdn.parleys.com/p/5148922a0364bc17fc56c60f/GarbageCol...
(* ) http://cdn.parleys.com/p/5148922a0364bc17fc56c60f/GarbageCol...
For dl.google.com, the first user of groupcache, it sounds like the Go runtime was working with relatively few, relatively large chunks of memory (think 2MB chunks of a Chrome binary). Since that doesn't have so many pointer-containing objects, just big chunks of inert data that don't need scanned, it's easyish. On the other hand, something like throwaway2424's case could be a GC stress test--many gigs of RAM holding a network of small, pointer-filled objects.
Really, the tl;dr may be to prototype/measure if you need to know how a GC will do. Assumptions are easy; data is hard.
And as for measuring... it's not cheap (developer time wise) to build a realistic enough model of your application and put realistic enough loads on it for long periods of time to test out each of your hypotheses (and GC theories do require long periods of load). So you have to narrow the space of choices informed by past experiences and sometimes by gathering semi-reliable folk lore. I'm currently engaging in the latter activity :-)
Atom also says 1.0's was more conservative, but, as Brad also said, still didn't scan "objects such as []byte" (meaning all plain-old-data arrays? who knows). The Go 1.1 Release Notes mention the collector becoming more precise, which was a particular issue on 32-bit because big heaps could span a lot of the address space.
You can see the GC source itself doing some per-type switching: https://code.google.com/p/go/source/browse/src/pkg/runtime/m...
At some point, this sort of discussion probably gets you less useful info per unit effort than just playing with a Go distribution, trying out whatever toy programs you find interesting.
And yes, unfair to ask someone to test every guess they have. But if you do want to know a bit more about what effect GC and Go would have, experimenting with toy programs isn't a bad way to get a feel.
I'm quite interested in the two mitigations you mentioned. Especially the "manage memory on your own option" ? I was not aware that there was a go language (that is, not a C extension) blessed way to do such a thing.
As for "reduce/avoid" creating garbage, how can I achieve that for a large heap caching application (like the one we're talking about here) ?
Some careful thoughts about the memory layout of your structs, especially which of them should be embedded and / or passed around by pointers and which of them shouldn't, might also pay off.
Another common optimization is to put the allocated objects back to a memory pool for later use. Take a look at the bufCache channel [1] from the bufio package for example (the http and the json package are using the same trick).
Another sometimes-useful way to save allocations--not for a cache, but in general--is just to turn short-lived allocations into longer-lived ones: if fooBuf is used and thrown away by each of several calls to obj.baz(), then make a single fooBuf and store it as a (private) field on obj, assuming that doesn't present thread-safety or other problems in the specific context.
There are certainly things you wouldn't use Go for, but now you know more about reducing memory pressure in GC'd languages. :)
The problem is, none of those things help when you have a large cache of small(ish) objects that need to be scanned fully at every GC cycle. There's nothing worse for your performance than stopping the world for a few seconds and using that time to wipe your L1/2/3 caches with completely useless data. That's the reason I called large-heaped caches the anti-pattern for a GC'd system.
Brad's answer about using large byte array allocations (and presumably, managing the smaller chunks manually) is a fair enough answer to this question. But of course in this case you'll be writing code to manage small allocations out of that big buffer yourself, thus negating the whole point of GC. But it might be a decent enough compromise if you like the rest of the language a lot (which Brad clearly does :-)
If you're worried about this, stick with C++.
The situation you describe, as you describe it, is (no longer - was it ever?) a pathological case. I run stuff with 12-20GB active objects and a few dozen http hits a second and the only times I get a latency over 100ms is when I reload/refresh some data from disk (1.5GB of gob), then it goes to 150-250ms typically.
It does not matter how sophisticated GC code is when you have millions of tiny objects from from crappy Java code with useless copying and referencing. This is why it uses so much memory.
It is very difficult to find any worse thing that Java in terms of wasting of memory. Take any Hadoop none and measure the ratio between memory used by Java processes and amount of actual data stored in memory. Near factor of 2?
Or, you know, you can avoid writing "crappy" code. Whereas with Go's GC, even proper code will have problems.
You are arguing something that is beyond discussion: Java's GC is better than Go's.
Even in Clojure code, for example, they implement a node of a binary tree as a hash-table which is a meaningless waste. Position based selectors would be good enough.
Java's GC is better than Go's - yes, in theory, on paper. In practice well-written Go code could outperform typical Java code, even with less sophisticated GC.
Of course, I cannot prove that, but there are some intuitions to support such claims.
It is true that you should describe Go as a garbage-collected language, but unlike a lot of other such languages, Go has C-like mutable arrays readily available. In practice it's more like a sort of hybrid, in that even though it's garbage collected you have a lot of opportunities to write code that still doesn't really use it, without having to "drop down to" C or something. I still hope to see some improvements in it before I could make a big commitment to it, but I've certainly got my eye on it.
Nothing stops you, except: - The determination NOT to do something the computer can do - Previous exposure to a proper, modern GC - The fact that you used Go to get away from this kind of shit in the first place
If you don't want to do that, don't. I'd never start the prototype with that functionality, for instance. But when you discover that you need it, Go permits it in a way that Python or Perl do not. Go is GC'ed and you are free to use that to the extent you want, but it's easier to escape from the GC than it is in many GC'ed languages. It may not be obvious from a casual reading of the spec, but there's a lot of ways around garbage in Go. Not like, say, Rust, and certainly not the same level of Raw Unstoppable Power as C++ (at a corresponding huge complexity cost), but it's a very interesting middle ground.
Note that if you were writing groupcache in Java, you'd probably end up doing the same thing and just allocating an expanse of bytes which you manage yourself.
I've often heard people say things that imply that "proper" GC's don't behave like whatever badness I'm experiencing currently. Can you give me an example of a Garbage collector that I could safely use for a large heap of small cached (ie. long lifetimes) objects without a significant throughput penalty and with deterministic latencies (ie., no random full system pauses) over, say, a tcmalloc based hand-allocated/free'd system ?