Go 1.4+ Garbage Collection Plan and Roadmap
golang.org
golang.org
I want badly to fall in love with Go, I've enjoyed using it for some of my projects but I've seen serious challenges with more then four cores at very low latency (talking 9ms network RTT) and high QPS.. I cheated in one application and pinned the go process to each cpu and lied to Go and told it there was only one CPU but then you loose all the cool features of channels etc.. That helped some but a Nginx/LuaJIT implementation of the same solution still crushed it on the same box, identical workload.
It would be nice if we have to have STW to have it configurable, in some environments swapping memory for latency is fine and should be configurable.
The way Zing handles GC for Java acceleration is quite brilliant, not sure how much of that is open technology, but it would be cool to see the Go team reviewing what was learned in the process of maturing Java for low-latency high qps systems.
channels work just fine in a single-cpu context. concurrency is not parallelism.
At least not without using one of the various netchan-alikes or rolling your own channel<->ipc solution, neither of which is likely to result in an elegant solution (note how the original netchan was abandoned by the Go team).
Doesn't http://godoc.org/runtime#LockOSThread let you do that?
Depending on your usage profile, there is very little more efficient from a throughput standpoint than STW collection, except maybe manual memory management with clever use of pool allocation. Even malloc()/free() will be slower in situations where you're creating a lot of short-lived garbage.
FWIW, no one serious about allocations in native code uses naive malloc()/free(). My favorite trick is in game programming where you have a current frame pool that you just reset every frame.
With manual memory management you have to worry about the ownership convention of each chunk of code, with GC you have to worry about architecting strong/weak references so as not to inadvertently retain everything (not to mention latency issues), and with reference counting you have to worry about ownership cycles. In practice I haven't come up with a better strategy than enforcing some sort of top-down hierarchy which effectively smashes most of the differences in cognitive load between manual/refcounting/GC. GC has slightly less upfront busywork but in practice the tooling tends to be poor so it's a wash (if that). In GC and refcounting it's easy for one inexperienced/tired/sloppy individual to create a massive leak completely out of proportion to the footprint of their immediate code.
Pools, in contrast, allow the same top-down approach with the very significant benefits that I don't have to think about the hierarchy at a finer level than the pool itself (which I have to do for the 3 other approaches) and that memory management mistakes don't typically lead to the globally-connected-component of the dependency graph sticking around indefinitely.
Another problem with pools is that you can't deallocate individual objects inside a pool. This is bad for long-lived, highly mutable data stores (think constantly mutating DOMs, in which objects appear and disappear all the time, or what have you).
Also, Go doesn't scan arrays that are marked as containing no pointers, so representing an index as a massive array of values has proven quite effective for me.
I'm not sure how much of it applies to Go, but I'm guessing it's quite a bit more complicated than the current collector. It also may only work on 64 bit X86 hardware (which is fine by me, but looks like a subset of what go currently supports).
There's a simpler overview in http://www.azulsystems.com/sites/default/files//images/wp_pg...
https://groups.google.com/d/msg/golang-dev/GvA0DaCI2BU/1EpYa...
The good news is Go has been chipping away at that stuff.
(a) A mutator may be stopped for up to 10 ms of every 50 ms by the STW collector.
(b) If more collection is needed, it will transition to a concurrent collector for the remaining 40 ms. Concurrent may pause stop a mutator for up to an additional 1 ms.
Therefore the worst case is 10 + 1 = 11 ms, not 1 ms.
So in terms of real time, your deadlines could be no less than 11 ms.
That's very interesting because LuaJIT has GC as well. Could you please reveal a bit more about the type of application, number of lines, and if the code is public?
Static compilation can do it to some degree as well, but in most mutation-happy languages (Go included) it cannot prove the allocation does not escape frequently enough to matter.
Allocation sinking is much more general. It works even if the allocation escapes (the allocation is sunk to the point where it escapes).
Additionally, all allocations can be sunk, not just stuff like the point class used in the example on that page. Growing dynamic buffers, string concatenation, etc, are all sinkable.
My understanding is that it's fairly tricky to implement (requires special consideration in runtime design) and so is not very widely used. Last I checked, the Dart compiler is the only other place I've heard of it being used.
FWIW its not really a thing Go would need anyway, given that you actually have control over your allocations.
If your primary motivation for using Go is lightweight concurrency, it has been developed for Lua, though in fairness not as extensively as goroutines:
http://github.com/askyrme/luaproc
http://www.jucs.org/jucs_14_21/exploring_lua_for_concurrent/...
That's why people are asking for pauseless.
60fps was fine pre-VR, but now you want to do 90fps times 2 eyes.
That's a lot of rendering.
If you want to update some physics model at 60 FPS, you will be updating every 16 ms. An occassional 10 ms pause is not going to prevent that. Anyway, any reasonable game design needs to deal with pauses that are caused by network congestion intelligently, which can often be longer than 16 ms anyway.
I don't even know what you are talking about wrt network congestion. What are you talking about??
I developed games for Android at one point. GC pauses were common there. In the 1.0 version of Android, the pauses could be up to 200 ms. Now, it's more like 5 or 6 ms on average, as they improved the Dalvik VM. The GC hasn't stopped a lot of good games from being written for Android.
People need to be aware of the realities of scheduling on commodity computer hardware. A 10 ms guarantee is actually very good. You can't get much better than 10 ms timing on commodity PC hardware anyway. Even a single hard disk seek or SSD garbage collection event will probably block your thread for more than 10 ms.
Whether this guarantee is good enough for you depends on what you're doing. But I guarantee you that 99% of the crowd here doesn't know what you need to do to get better guarantees than 10 ms anyway (hint: it involves a custom kernel, not doing any disk I/O, and realtime scheduling).
But that is totally different than the GC situation. With a STW GC you cannot even run, how do you expect to be able to display anything on the screen? Even with a non-STW GC, the reason the GC has to collect is because you are running out of memory (unless you are massively over-provisioned), and if you are out of memory for the moment how are you going to compute things in order to put stuff on the screen?
Accessing disk/network/etc induces latency, yes, but that is why you write your program to be asynchronous to those operations! But this is a totally different case than with GC. To be totally asynchronous to GC, you would need to be asynchronous to your own memory accesses, which is a logical impossibility. I do not see how you even remotely think you can get away with drawing an analogy between these two situations.
Having the latency be capped at 10 ms is an extremely powerful guarantee. It means that even if your machine is not fast enough to play game X one year, next year's machine will be. Because the GC latency doesn't change, and the rest of the program logic will run 2x as fast because of more memory and cores.
If you want to act shocked, shocked that GC has an overhead, then go ahead. If you want to pretend that nobody can ever build a game in a GC langauge (despite the fact that hundreds of thousands have, on Android and other platforms), then go ahead. Hell, even if you want to continue to use C or C++ to squeeze out every last drop of performance, then go ahead. But I find these comments really disingenuous.
Yes, Java claims to have an efficient GC (on paper). In reality, everyone know how it works (it bloats, and GC pauses are increasing dramatically).
On the contrary, Haskell has less sophisticated GC but much smaller and more predictable GC timings.
btw, some really smart people are suggesting that it is the code organization, not GC algorithm is a crucial factor. http://home.pipeline.com/~hbaker1/LinearLisp.html
Go is better suited than Java for those kinds of applications, because it's easier to avoid allocations (Java doesn't have stack allocations, but Go does.) Also hard upper limits on GC time are very helpful for those cases where allocations can't be reduced any further. The standard library additionally has a Pool type that allows for reducing pressure on the GC through object reuse.
Go also lets you have fields in place in structs, like C and C++. In Java, all fields (except primitives) are references.
It's fine for basic stuff but if you want to really listen to full RTB bid stream and have real time responses you need pauseless GC to win. Or you have to budget for considerably less time to make a decision.
Also there are places in RTB where you have 10ms RTT and Go fails ought right in that situation at scale. :-)
type s struct {
x int32
y int32
}
v := make([]s, 1e5)
The object v uses 1e5 * 8 bytes of storage. Equivalent Java code would require boxing of each s value. That is, instead of []s, you end up with []*s and allocating each s.http://www.chrisseaton.com/rubytruffle/pushing-pixels/
Graal can do normal escape analysis, and so scalar replacements, but it can also do partial escape analysis - that is doing scalar replacement even if an object will later escape.
It has a pauseless generational collector called C4 (which is detailed in the Garbage Collection Handbook). Gil Tene, the CTO and founder of Azul, appeared on the Go mailing list a while back and outlined the steps towards implementing the C4 collector in Go.
https://groups.google.com/forum/#!topic/golang-dev/GvA0DaCI2...
This looks like progress towards that. But I don't know if C4 is any kind of concrete (or even speculative) goal. I would love it if standard Go was running a C4 GC implementation.
Yes, but 10ms is too long to be useful for games. I'd rather take 1ms every frame than 10ms sometimes.
And, as I mention whenever this comes up, Minecraft is written in Java and suffers from huge GC pauses, yet it's one of the most successful games of the last decade.
To return to the topic at hand: 60fps is 16ms per frame. The proposed GC would achieve sub-10ms pauses. So, worst case you might drop the odd frame. To assume that Go is unsuitable for games because of garbage collection is to ignore reality.
Furthermore with games you can do tricks like allocating out of an arena with bump pointers and then resetting the pointer at the end of the frame.
Lots of games are developed with JVM, .NET, or Lua. GC doesn't seem to be a show stopper for them, you just have to be smart about allocations.
And yes, it's a problem.
See the other comments here about VR. With VR you want to render at 90 frames per second, in other words, you get 11 milliseconds to draw the scene twice (once for each eye). That is 5.5 milliseconds to draw the scene. If you pause and miss the frame deadline, it induces nausea in the user.
But this comment drives me up the wall:
"GC doesn't seem to be a show stopper for them, you just have to be smart about allocations..."
The whole point of GC is to create a situation where you don't have to think about allocations! If you have to think about allocations, GC has failed to do its job. This is obvious, yet there are all these people walking around with some kind of GC Stockholm Syndrome.
So now you are trapped in a situation where not only do you have to think about allocations, and optimize them down, etc, etc, but you have also lost the low-level control you get in a non-GC'd language, and have given up the ability to deliver a solid experience.
Bad trade.
Nope. The point of GC is memory safety.
GC also means you don't have to think about freeing memory, which is important in concurrent systems.
But even if GC was about "not thinking about allocations", what's bad about only having to think about allocations when it's important? Code clarity trumps performance, except at bottlenecks.
You can get memory safety without GC, and a number of GC'd systems do not provide memory safety.
If you think that, for concurrent systems, it is a good idea to let deallocations pile up until some future time at which a lot of work has to be done to rediscover them, and during this discovery process ALL THREADS ARE FROZEN, then you have a different definition of concurrency than I do. Or something.
If you want to know about code clarity, then understanding what your program does with memory, and expressing that clearly in code rather than trying to sweep it under the rug, is really good for clarity. Try it sometime.
I would say for most uses, that's true. For high performance code, it's not. You have to be aware of your allocations on the fast path whether you are using malloc, GC, or something else. Schemes used in C++ to reduce allocations with malloc (which is slow as dirt) also work in Go.
Let's get real here for a second. A maximum pause of 10ms is one-third of your frame budget at 33fps. If you reduce the garbage you're generating it won't even be that much. I don't see how that makes it a show stopper for games.
At the same time, I don't see how it's a great language for games either. Go is slow at calling into C, like 2000 cycles slow. And to do any kind of work with with the cpu vector units, you need to call into C. That seems to be a bigger issue than the GC in my opinion. Rust seems much better suited for games, and I would expect to see some triple A titles using the language once it's more mature.
No, Lua's current GC is better for games. It is an incremental GC. For example World of Warcraft originally had performance issues caused by the standard STW GC in Lua 5.0 but those largely went away after they switched to Lua 5.1 which was the first release featuring the new incremental GC.
Anyway, thanks for seconding their recommendation of the book; I'm looking forward to reading it -- and glad I saved the 35 clams!
[1] https://www.evernote.com/shard/s249/sh/53d7b93a-d737-41c4-8e...
[2] https://www.evernote.com/shard/s249/sh/a1c55244-d64e-40cc-a6...
http://camelcamelcamel.com/The-Garbage-Collection-Handbook-M...
"Hardware provisioning should allow for in-memory heap sizes twice as large as reachable memory."
So I know "memory is cheap"(TM) but surely a 100% physical RAM overhead for your memory management scheme is worth at least a small amount of hand-wringing. No?
See for an example:
http://sealedabstract.com/wp-content/uploads/2013/05/Screen-...
(quoted in http://sealedabstract.com/rants/why-mobile-web-apps-are-slow...)
According to the graph, most GCs start being really slow even with 3 times more memory than needed with the manual management.
First of all, the "most GCs" are the non-generational GCs.
That's well known. The amount of tracing work that a non-generational allocator has to do per collection is proportional to the size of the live set. Thus, collecting less frequently (by increasing the heap size before another collection has to occur) makes GC faster, using time roughly inversely proportional to how much bigger you make the heap.
Generational collection can greatly mitigate tracing of the live set for when you have many short-lived objects that never leave the nursery.
Second, the benchmark compares the allocation/collection work of various garbage collectors vs. an oracular memory manager using malloc()/free(). Not only do alternative solutions that don't use automatic memory management not necessarily match this performance (naive reference counting tends to be even slower, pool allocation also can have considerable memory overhead, etc.): more importantly, it's an overhead that applies only to the allocation/collection/deallocation work. For example, if your mutator uses 90% of the CPU with an oracular memory manager, then a 100% GC overhead means 10% total application overhead.
Correct me if I'm wrong, but even the "generational" GCs present their own problems: the "costs" of having the GC increase not only with having "too little" memory but also with trying to use "too much" memory (as "more than e.g. 6-8 GB, which can be needed on the server applications). As far as I know, only Azul's proprietary GC is claimed to avoid most of the problems typical for practically all other known GCs. fmstephe in his comment here linked to one discussion where the Azul's GC author participated. But I nowhere read the claim that any GC doesn't need significantly more RAM than manual management.
Simply look at the picture you referenced and fully understand what it says. Read the accompanying paper also.
> Correct me if I'm wrong, but even the "generational" GCs present their own problems
Every memory management scheme has pros and cons, yes.
This is not what you wrote. You said that "most GCs start being really slow even with 3 times more memory than needed with the manual management", while the generational mark-sweep collector has essentially zero overhead with 3x RAM in that benchmark. The "most GCs" you're referring to are algorithms that are decades behind the state of the art.
Also, "really slow" is a fuzzy term and I am not sure how you come to that conclusion from the image.
Remember, they're compared to an oracular allocator that has perfect knowledge about lifetime/reachability without actually having to calculate it. That ideal situation rarely obtains in the real world. The paper uses this case to have a baseline for quantitative comparison (similar to how in some situations speeds are expressed as a fraction of c), not because it represents an actual and realistic implementation.
Your only arguments: after showing that I wrote "most need even 3 times more" then you give an example of one which needs 2 times more. Then you complain that "really slow" is fuzzy. Then you claim that "ideal situation rarely obtains in the real world."
I asked you for the graphs and links.
If you're struggling with understanding the paper, there's really nothing more I can do to help.
It's difficult to measure this overhead, because you design the program around the limitations of your language. The point is that, by rejecting GC, you gain certain flexibilities that can enable higher-performance designs. Merely removing GC from a Java app doesn't capture that.
For a lot of applications on a GCed runtime 100% isn't an unreasonable number though. To get the overhead down you can manage your memory more carefully (pooling, keeping most objects pointerless) and tell the GC to run earlier.
https://gist.github.com/jboner/2841832
If you trade cpu usage for heap size, you'll lose in all modern cpus. In fact it's getting to the point where worse algorithms that are space-efficient are starting to beat the best known algorithms in practice, even if their constants are different. Dual TLB lookup when running virtualized doesn't help with that either.
If you can avoid a memory read with 4-500 assembly instructions (say a C/C++ function of half a screen), that is actually worth it. I did a test on my newest machine that showed that it's actually more efficient to run the sieve of Eratosthenes algorithm for prime numbers under ~800, than to have a table in memory with those primes.
It also means that you should easily be able to execute about a page of inlined C/C++ code in the time it takes the cpu to simply jump to a virtual function. Go and Java only have virtual functions.
Likewise, having programs work directly in compressed memory is actually a net-win. E.g. storing protobufs in memory and unpack single fields out of them when used instead of using unpacked structs directly is actually a win for even quite large structs. Not for megabyte-long ones, but you know, in 2 years there'll be a new intel micro-arch and who knows ?
Good luck keeping objects pointerless in Java. Or go for that matter (you can't use any interfaces or pointers if you do).
But don't take my word for this. Here's a trivially simple test. Gcc can compile java programs down to machine code. The memory model of those compiled things is different from the jvm. Take what you consider a slow java program, gcc it, and execute it. Note the 100%+ improvement in speed. (only really works for memory intensive stuff, of course). Rewriting it in C++ (assuming you know what you're doing) will get you another 100%+ improvement.
When it comes to designing the data structures your app uses you have a point but I think it's not directly to your app or runtime's policy for reclaiming freed objects. This does raise an idea that hadn't occurred to me before though: once you have a generational moving collector you may actually get virtuous cache effects because your nursery is densely packed with live objects.
...
Ok, it turns out that this has not only been studied a little but but at least PyPy tried this at one point (I looked at the code but couldn't find it).
http://books.google.com/books?id=zVbbkWnxDP8C&lpg=PA99&ots=1...
http://pypy.readthedocs.org/en/latest/garbage_collection.htm...
It is a shame too. I love writing go
There are lots of other contexts where the performance profile outlined in this document are sufficient. If you love writing Go, then I'd suggest staying out of the financial industry and robotics.
And avionics/aerospace.
And self-driving cars. And medical equipment.
etc, etc. You can list lots of fields for which this is unacceptable, and they are a lot of the really interesting fields.
As far as gc bits, the Azul Systems pauseless gc is certainly as state of the art as it gets. Too bad it is for java
Like it or hate it, electronic trading is the natural progression of the markets, just like the car eclipsed the horse. Low latency in trading applications affects me when I move my 401k as much as it affects the big electronic trading firms. Thinking it doesn't affect you is a mistake :)
But for 90% of those trades, it's a block limit order that sits on the order book until it's filled.
For many trades this is irrelevant. For those where it does matter, it is the very last, smallest part of the system -- the execution engine -- where timing is everything. That is the part of the system where HFT engineers are building hyper-speed interconnects and optimizing every single instruction, but it is not relevant to the majority of the stack.
Delays everywhere else in the system are largely meaningless -- yeah prices may have changed slightly, but unless you've mastered market timing (which no one has), it will generally even out and is irrelevant.
Sure in a small percentage of places, maybe, but show me a real study where this is shown to be a real factor in deciding a successful company.
For games, 10ms is way too high. I'm working on Oculus Rift stuff at the moment and really, you have about 13ms per frame to play with in total, and skipped frames are bad.
However, for typical web apps without low-latency realtime requirements, a guarantee that the GC will not add more than 10ms latency at once, and no more than 20% overhead (in the same thread) seems pretty good, especially considering that it can do this under high load conditions with relatively constrained resources (Java is probably faster, but likely to use more memory; most other languages will be significantly slower).
The one outlier report I mentioned was of GC randomly stopping processing for several minutes, which clearly isn't within the bounds of acceptable response time for web apps. This will completely eliminate the possibility of that happening, which is good.
All in all it seems like a decent trade-off for the sorts of applications Go is already fairly good at. The stuff it's no good at might come later, but in the meantime everyone knows you can't use Go for those things (in fact you can, if you can design your entire application to use stack allocation only, in which case you can turn off the GC entirely, but that's easier said than done).
> The one outlier report I mentioned was of GC randomly stopping processing for several minutes, which clearly isn't within the bounds of acceptable response time for web apps
That's not acceptable for any kind of app and the problem gets bigger the more heap memory you have. The sweat spot for GC enabled apps seems to be 4 GB - anything bigger than that and one should prepare for surprises.
Does Java give any guarantees about garbage collection latency?
The LMAX Disruptor inter-thread communication library[1], that is the heart of the high performance[2][3] limit order book used by the LMAX Exchange is written in Java.
Just because a language can't guarantee that the duration of GC stops are below a certain threshold, does not mean that you cannot write programs in this language that offer better latency 99.99% of the time.
[1] https://lmax-exchange.github.io/disruptor/
Java has multiple JVMs and garbage collectors available.
The concurrent GC (CMS) in Oracle's JVM doesn't give any guarantee. The new G1 GC in Oracle's Java 7 does provide guarantees while trying to minimize STW pauses. Azul Systems' Pauseless GC is advertised as being completely pauseless, though it is commercial.
Quite painful is that for small heaps, the more you try to do concurrent garbage collection with small latencies, the more throughput suffers. On the other hand, the more you add memory, the bigger the latency. So depending on the app, the memory layout and its access patterns and the hardware used for deployment, one has to pick a GC strategy, as there's no one size fits all.
I guess my overall point is that the GC guarantee in question is an upper bound. It's perfectly acceptable for a financial application to only deliver x ms latency 99.9% of the time.
Also, the JVM is only one component in the stack. If you aren't using a real-time OS kernel, you're not guaranteed any maximum latency anyway.
Somehow very niche HFT trading platforms have come to define finance, when the reality is that the massive bulk of the industry runs on Excel spreadsheets, Java, or Fortran. An extraordinarily small subset of the stack has real-time needs, and is engineered as such.
The financial industry is a huge industry with an enormous variance of needs, and the overwhelming bulk of applications have absolutely no such real-time requirement.
I should mention that I've built systems in the financial industry for about a decade now, so I always marvel when people make broad, untrue, absolute statements.
I want to like GO but a language that targets native development but still uses garbage collection just seems like an odd pairing to me. Maybe it's just me especially since I rarely get an opportunity to do native development.
RAII is but one resource-management strategy, and one intrinsically linked to C++'s memory model. But if that's what you're looking for, Rust uses the same: http://doc.rust-lang.org/std/ops/trait.Drop.html
And in the absence of cycles, refcounted implementations have similar properties (although of course that assumption will break with alternate non-refcounted implementations if refcounting is not part of the language semantics).
> I want to like GO but a language that targets native development but still uses garbage collection just seems like an odd pairing to me.
Not that rare. OCaml is native and GC'd, so is D, or most lisps with native compilers. Hell, Boehm's GC exists.
http://nimrod-lang.org/manual.html#reference-and-pointer-typ...
The price you pay is you have to manually avoid cycles by annotating some references as weak.
It's a very nice scheme, actually.
Correct, Swift is built on top of the obj-c runtime (more or less) and also uses ARC.
Actually, even without cycles the destruction time is in general unbounded. Just think what happens when you allocate a very long linked list one element at a time and then drop the head. With a bit of ill luck or intention, you can make each element be allocated from a different page with enough different pages so that they fall out of the TLB. In that situation, even without having pages written out to disk, you can expect each element to take ~ a thousand cycles to free. On a million element list, that's ~500ms for freeing the head.
The problem is, compiler magic gives the illusion of safety, and then you segfault and remember that you're in a non-memory safe language. Swift is a big improvement, but it's still basically a DSL on top of objC, with all of the downsides that implies.
The segfault I hit yesterday had to do with a CoreData API that didn't retained its argument, when I assumed it did. The bug doesn't appear when compiling with -Onone, but does appear in -O. Compiler flags are another place where Swift is still very much in C-land, rather than Java/Python/Ruby land. You still need to understand the C compilation model, and e.g. getting your ARC flags wrong in a cocoapod will still cause crashes.
To be clear, swift is a big improvement over objC, but you can't pretend it's a VM managed language.
Memory safety does not require VM management (or so Rust is trying to prove).
I agree, with a few qualifications. I think Haskell has already proved it's possible.
I think Swift + ARC isn't a strong enough guarantee of memory safety for two reasons:
1) Swift has to interact with a huge amount of objC code, of varying degrees of quality (some Apple code, some OSS code via cocoapods). Swift itself would be much safer if all objC was rewritten in Swift.
2) ARC isn't as strong a guarantee as either GC or Rust's pointer semantics. For example, it's easy to hold weak references to something you thought was retained, and then dereference an invalid pointer.
So yes, you can have a memory safe language without a VM, but I think you need GC and/or much stronger pointer semantics, ala rust.
Haskell has a GC, no?
Why? I don't see what is weird about that combination. Garbage collection is usually something that is used in high-level languages, and high-level languages might as well be natively compiled as opposed to interpreted, running on a vm, etc.
If anything, C/++ seems like the odd pair out with their more manual memory management among native languages.
Catching up old GCs via concurrent GC states is fine and tandy, but it is still just catching up, and it requires GC state. Cheney not. And a typical Cheney GC is 3-10ms not 10-50ms.
Garbage collection is required but even a hybrid STW only reduces latency but doesn't eliminate it. Nor is there seemingly any foreseeable way of allowing developers to issue a GC request in a timely fashion.
What if, prior to enacting GC, Go concurrently shifted from it's current memory allocation to an exact clone, cleaned up, then shifted back? Or maybe it cleans as it's cloning, enabling a shift to a newer, cleaner allocation? Sure, there would be latency during the switch, but it would be considerably less than stopping everything and waiting for GC to finish.
[1] https://en.wikipedia.org/wiki/Mark_and_sweep#Moving_vs._non-...
I guess then the question becomes why would the roadmap avoid such an implementation?
http://www.azulsystems.com/sites/default/files/images/c4_pap...
All in all, depends whether you're serving for a turn-by-turn roguelike, an RTS or an FPS. Won't be an issue for the first one, may not be for the second one, the latter though...
So there's no reason why your typical indie style game would not be written in Go, but of course if you want to write AAA style 'push the boundaries of realtime graphics and physics', Go or any other garbage collected language is not a very likely candidate.
You can deal with deallocation chains in reference-counted systems without having arbitrarily-long pauses. You keep a stack of <pointers to items to be removed + current child index in item>. When something's reference count goes to 0, push it onto the stack, with a child index of zero. Every so often, pop the last item off of the stack, if (the child index is after the last child) {deallocate the memory}, else {decrement the reference count of the child of the object indexed by the child index, pushing that onto the stack with a child index of 0 if the child is now dead, and push it back onto the stack with an incremented child index}. You can do any/all of this in batches.
It doesn't, but it most real-world implementations of RC it is.
The point, (again note I'm not an expert in this stuff) is that there's no such thing as an infinitely responsive real-time system - even hardwired analogue electronics, or a program running entirely in CPU cache, have a response time. So realtime always means that the system responds acceptably quickly for the application, without slowing down.
Now as to how this applies to games, obviously a 10 ms pause is not a great help for updating a 60 Hz game every 16 ms. But for some lower speeds it could reduce latency and give an overall smoother feel to networked gaming. A couple of other posters on this thread have already mentioned tricks to get around this, if multiple processes can share the load and gracefully garbage-collect.
But on the whole it seems like offering some configurability would make Go more versatile. In the extreme you might want to turn off the GC and make use of a manual free-type function. (I gather the Go developers oppose this on philosophical grounds).
When GC comes up here (and not only on Go things), a lot of comments pop up that are like "welp, [tool] is no use for me if GC is involved," sometimes with a horror story or worst-case scenario. Sometimes I'm sure that's entirely rational and comes down to the type of app they work on (I don't wanna rewrite Unreal Engine in Go), but sometimes I wonder if whoever's saying it just doesn't know much about GC in practice.
"Go, the language with zero nines of availability."
Even erlang will have some pause times, that has nothing to do with availability. And as mentioned, this is only when GC pauses happen.