Large-scale, semi-automated Go GC tuning
eng.uber.com
eng.uber.com
https://eng.uber.com/jvm-tuning-garbage-collection/
Here's another blog from Uber on JVM tuning.
Some notes:
1. They have options
2. They have a powerful gc log, no need to roll your own finalizer thing like this
3. Choosing the max heap size, which seems like what they actually wanted here, is trivial
You could say "but that's more complex", but to me it's that the JVM has far more mature features, tons of tooling and options that they could explore and adopt, and that the obvious wins are trivial to achieve through basic parameter tuning.
Further, this GOGC parameter seems to be a very weird knob. The knobs you'd run into with the JVM will often be a lot more straightforward ie: a static value for heap size vs some number that's based on a working set percentage.
I wonder if over time Go will end up with a tunable GC.
The "debate" is usually between people who have committed to either path (creating tools or learning tools) and think that just because the marginal cost of my approach, for me is 0, it must be 0 for everyone else. Which is patently false.
With Go there's one parameter, and in my opinion it's a very strange one. It also seems strange to have to (imo) hack GC metrics in using finalizers, whereas with the JVM it's simply provided to you.
Full disclosure though, I think Go is a bad language, so I'm biased.
It’s weirdly black and white to assume that just because someone thinks a language is “bad”, that opinion doesn’t have nuance.
Full disclosure, I also think go is a bad language ¯\_(ツ)_/¯
Hell, I enjoy writing C but I still think it’s a bad language.
Nothing I’ve replied here has been an “attack” on you. I simply tried—gently at first—to suggest that “x is bad” should not be equated with “x is irredeemable” which is not exactly a charitable or reasonable interpretation in the context in which that statement was originally written.
Further I directly expressed a more nuanced opinion as an example in my and you still chose to discard that nuance and interpret the opinion as black and white.
I think you’ll find that on a scale of positivity from -1.0 to +1.0, most people perceive “bad” as somewhere along the lines of “< 0.0” and not “= -1.0”.
I have zero interest in rehashing an argument that’s been made here hundreds if not thousands of times already, and by others far more eloquent and convincing than myself no less. Feel free to read my post history. Or simply search for virtually any golang-related post on this site. Whatever arguments you find, I probably agree with at least 80% of them.
Have a nice day.
I believe go is a bad language. I have nuanced, lengthy, and detailed opinions behind that belief which stem from 24 years of software engineering, 4 years of professional experience specifically with golang, and professional experience writing, deploying, and maintaining production software using C, C++, Rust, Java, Ruby, Perl, and JavaScript. And I have zero interest in rehashing the past twelve years' worth of arguments against golang with someone who's repeatedly signaled a frustrating level of obstinance.
Whatever wild conclusions you choose to jump to from there are your own doing, not mine.
Also: I don’t believe you. You have provided no evidence that you actually have a nuanced opinion, you’ve simply insisted upon it it’s possibility. And I don’t think there’s any reason I should believe you.
It feels like trying to get trumps tax returns. “They’re great returns” he insists, but he will generate all sorts of arguments to try and stop you from actually seeing them.
I could talk a lot about why I think it's a bad language, it would be hard to summarize it since I'd want to cite Pike's talks on "simplicity", articles on Go's GC implementation, discuss error handling, what I think makes a language "good", etc.
Example transpiler input / output: https://github.com/nikki93/gx/blob/master/example/main.gx.go... becomes https://gist.github.com/nikki93/97ff376abb6718427387bb9cca2f... Can call to C/C++ (including templates) w/o overhead.
That said, for logic that is heavy on async and escaping closures like how a lot of Go server code tends to be, a GC is maybe a reasonable tradeoff?
I was using C++ for data-oriented gameplay coding, and I was interested in exploring making a language frontend that compiles to it to clean it up, as a side project (lots of dark corners to run into with C++, and I collected some experience on what those were since I'm managing a C++ game engine codebase at work). I needed a core that was basically a cleaned up C, which is what the C-ish core of Go is (the part other than goroutines, channels and GC), and Go has a good parser and typechecker library you can use. Goroutine and channel are cool for distributed server code or whatever, but not actually that useful for game programming. The main thing is having structs, procedures, some nice ergonomics over those (slices, type inference, non-escaping lambdas, occasional generics) and then metaprogramming so you can reflect over the data structures and have serialization and inspector UI. These are the elements actually relevant to game programming.
Why didn't you consider D in 'Better C' mode? (https://dlang.org/spec/betterc.html) Not only does it already exist but with very high likelihood is more polished (by virtue of the man-years already invested into D) than a single person's ad-hoc compiler of a subset of Go to C++ could probably be. Unless of course you absolutely needed to use some pre-existing Go code...
Just because something has a bunch of years in it doesn't mean it's a good idea for a specific context. The transpiler I have here is just 1500 lines of code and captures all the semantics it currently supports. It uses Go's parser and typechecker from the stdlib and feels on the whole more polished than D as a result (generics are definition-checked, Go's package / module system just work, all the existing Go editor support and godoc etc. just work, ...). It's much easier and straightforward to metaprogram by just editing this simple piece of logic than squeezing it into language features (I've also done the same engine in Nim, explored in Zig, and written it once over in C++). I can, for example, make it so if you mark a function a certain way, it's also compiled to GLSL and useable as a shader (with structs shared). Or make it so types marked a certain way have all their pointers reference counted. There's way more control in this scenario, and the point is to have control to take matters into one's own hands and actually improve things.
What is a "proper abstraction"?
Or feel free to manipulate the voltage in some wire, and make sure that it is reliably understood at the other side as the same bit pattern you sent, but I prefer issuing an HTTP packet. These are all abstractions, hell, there is no field building as much on abstractions as IT does. We have to be on like 8-9 levels of abstraction to even do anything non-trivial.
This is what the game code looks like: https://gist.github.com/nikki93/0425d9ead9eb7810075434d006f3... -- I don't think CSP helps much to improve on that while keeping serializability of state.
This is kind of a semantics question, I think. The vast majority of games need some sort of lifetime management, and it's just a question of who does the lifetime management and what mechanism you use for it - refcounting, a mark/sweep GC, freeing everything at certain points of time, arenas, etc. If you're using an entity/component system to manage lifetime you have a GC - you wrote it.
In my experience shipping games in C# compared to shipping them in C/C++ - you can do everything without the GC touching your stuff if you're really dedicated, but it's often not worth the trouble considering that any modern GC can handle scattered temporary per-frame allocations for you no problem with very minimal pause times, as long as you're thoughtful about it and the set of objects it needs to walk isn't too big. For example, if your data structures mix native data with pointers to object instances, a GC will have to sweep all of that data - splitting texture references out of a big table of draw calls means that the draw calls are now pure data and they don't need to be swept by a GC.
Personally I prefer always having access to a GC because it means code that doesn't need careful lifetime management can be simpler to write and doesn't have issues like double-frees hiding inside it - things like automated tests, configuration UIs, debug consoles, and things you run once at startup or when loading a level. You can often go back and optimize some of this stuff later, too - for example LINQ is a notoriously messy feature in .NET's standard library that allocates tons of short-lived garbage, but the compiler makes it possible to replace all those LINQ data structures with non-allocating ones without having to rewrite your queries - but doing that moves costs elsewhere.
If you're getting specifically harassed by pause times you're likely going to be paying costs with other systems, like if you use refcounting any time you touch that refcount you're burning cpu cycles and pushing other stuff out of cache (and the refcounting gets much more expensive if you have to use atomics for thread safety).
But yeah I've found that the entity component data structure is a good lifetime management system very well suited to the game scenario, so there's no need for a different / more complex thing. And this is an exploration in how that + a simple / ergonomic language around it (along with growable arrays (slices)) pan out when making games in practice. There's no manual free calls or lifetime management anywhere in the code.
Re: "using GC but then needing to / being thoughtful to make sure it's going ok" -- that's the thing. It seems better to not have to need to think about it, by having a system that's better suited to the thing you are working on. You also don't need to "often go back and optimize some of this stuff later too" because it just already has good performance with the straightforward code, and you don't need to add complexity. "splitting references out / pure data" -- that is indeed what the language nudges you to do by only having pure data. Essentially: yes, you can achieve the desired thing with intentionality and extra cognition in a different system (the same was true with other kinds of cognitive overhead in C++) and this is an exploration in developing a language + tools that focus on and bias toward the desired thing by default. Like I'm imagining folks getting started with gamedev using this + the integrated tooling and internalizing the practices you're talking about (that's a stretch vision, the current scope is to just build and test it in the context of one specific game project).
I'll be releasing this engine + an example game with it soon, but here's what the code for this main game project (the one in the video) looks like (all the components, the top-level game loop, and then some example game logic): https://gist.github.com/nikki93/0425d9ead9eb7810075434d006f3... It's just data structures, and then functions that do gameplay stuff on them. No lifetime management.
Mark and Sweep is antiquated for GCs. There are some edge cases where it is still useful. But most heavily used sytems should use more modern GC algorithms, e.g. generational GCs.
So there isn't a reason to use it and then work around it. If anything I think it's because the ergonomic language work after C++ (C#, scripting languages) tended to include GC so using them meant having it. This is an exploration in having a language that doesn't do that and preserves ergonomics (and it's working / promising).
All that said, the GC thing isn't the main or only reason that motivated this approach, it's just one of the points among everything else. Portability (the resulting C++ compiles and runs in Wasm, native desktop, mobile is supported) and control (being able to decide language and resulting execution semantics, having direct integration with tools) are the main things, at a high-level.
Second, we simply don't know, considering all aspects including human resources and engineering economics, whether Xerox, DEC, or IBM's alternatives would have fared better than UNIX. I'm open to learning why/how if this is a shut case in your opinion.
There's been a lot of back-and-forth on how to improve the tuning situation in go. See https://github.com/golang/go/issues/42430. It seems like Michael Knyszek will be doing something about it for go 1.19.
Ok, go has GOGC env. Ok, you can tune it based on stuff. Do they tune it live as part of process life or is it pre-comoputed at start? Is gogctuner a library?
> As we mentioned above, manual GOGC is not deterministic
What?
More importantly, why the engineering effort is spent on that tool as opposed to just trying to reduce allocations. I've spent countless hours trying to reduce allocations of the hot path. This is a good strategy - Go GC cost becomes negligible if it doesn't have anything to do!
But then it hit me. The missing context is probably other services/tenants interacting with system resources.
Am I wrong in reading it as: In time of the low load, they want to burn less CPU at the cost of more memory. At peak time, they do business as usual. Reducing GC frequency at low load, is generally meaningless. In most systems most operators care about performance/cpu/latency on peak.
Unless....
Unless you have other tenants. They probably run batch jobs at low times, and if that is the case, then indeed, burning CPU for low utilization GO jobs is a waste of CPU.
But otherwise agree with you.
uber was able to solve there issue so the GC in golang seems perfectly fine.
The optimizations the JVM does in the newer ones are well understood optimizations but don't address the root issues with implementing generational GC in golang.
the issues was with the write barriers and requirements around moving data which can't be avoided in generational GC implementations because you have to move data and update pointers.
most of the benefits of generational GC doesn't exist in golang because of escape analysis allocating data on the stack.
there isn't anything fundamentally wrong with golangs GC. attempting to apply the solutions for the JVM to golang is fundamentally flawed and ignorant.
the languages have fundamentally different approaches to memory allocations. and the reasons behind the JVM implementations simply do not exist in golang.
the new pacer is in the works to attempt addressing many of the edge cases. https://github.com/golang/proposal/blob/master/design/44167-...
you seem to lack fundamental understanding of what the actual issues in the golang runtime are with relation to its GC.
I'm not saying that Go should use GCs like Java's. I'm just saying that Go's GC does not work well enough for Go's needs, and that Java's GCs now deliver a better experience. Maybe Go needs something entirely different, but it does need something better than what it has now.
The biggest challenge with Go in production is that, as the article points out, Go doesn't have a maximum memory setting that can be used to tune the GC. If you run Go with a memory limit on Kubernetes, it's common for it to simply run out of memory rather than using the limit for backpressure.
This is kind of surprising, since Go is so entrenched in the Kubernetes world (and both Docker and Kubernetes are written in Go). It's possible that Google itself doesn't develop that much stuff in Go, and that the Go team doesn't have a huge incentive to innovate in this area.
There have been a few attempts at improving heap management, including an aborted attempt at respecting ulimit [1], a promising implementation of a SetMaxHeap() function [2], and a propsoal for dealing with backpressure [3], but these projects have mostly failed to get proper traction. It's a complex problem that needs a cohesive solution.
Fortunately, there is now a proposal [4], which has been accepted, to add a soft limit to Go, which has a more thought-through design [5], though I'm not sure if it's being actively worked on yet.
I'm also not sure if that proposal, when implemented, will make the Uber approach redundant, or if these are in fact complementary. If Uber could open-source their library, it might be a good solution until Go itself has better GC management.
[1] https://github.com/golang/go/issues/5049
[2] https://github.com/golang/go/issues/16843
[3] https://github.com/golang/go/issues/29696
[4] https://github.com/golang/go/issues/48409
[5] https://github.com/golang/proposal/blob/master/design/48409-...
A win is a win, and its still a very nice saving for barely touching the application code.
Discord had an issue with that for an LRU cache service, where their memory usage was basically constant (very little garbage generated), but because the heap was quite large the pacer would trigger a huge CPU spike every 2mn as it would need to traverse the entire thing, looking for something to release (which would not exist).
However, even though Discord's article is technically out of date in terms of their exact numbers, the principle still holds, just at larger scales. If one keeps scaling up, eventually one will encounter fairly fundamental and difficult problems that take odd solutions, and no fully automated memory solution will solve them.
I would observe, though, that these complaints are arising at a very significant scale. It is a common error in programmers to assess their needs as if they are going to be writing code running on a hundred servers maxed out on the resources at near 100%-CPU when in reality their code is going to comfortably run on one instance with 5% of one CPU in a day.
I say without hesitation that if someone is looking to run dozens of maxed-out servers, Go is a bad choice and it is a mistake to even start writing that code in Go. (There's many even worse choices; if Uber was trying to write the same service in Python or something... yeowch.) But if someone rejects Go because it can't hit that use case, but the use case couldn't possibly hit that scale unless every person on the planet become a customer five times over, that's making the exact same mistake. Go is a good solution for many very common use cases, but it's not that hard to do some Feynman estimations at the start of a project and notice that it's getting kind of close to the comfortable limits for Go.
(Even growth isn't really an excuse. Resources are so abundant that you should take a log-based view, or an exponential-based view if you prefer. I like to have an order-of-magnitude buffer minimum in my design for the largest possible scale I could face, and most of the time that's pretty practical nowadays. If I have a case where Go would work, but I'd only really have roughly a factor of 2x growth before it would become a problem, I wouldn't use it. It's too easy to consume that by either usage growth, or future changes in what the system needs to do, or error in the Feynman estimation. But resources are, as I said, so abundant that by the time I'm maxing out a 32-core or 64-core system with however much RAM that comes with nowadays, I'm running a lot of stuff.)
I would be curious if they've got a "rewrite in Rust" effort going. Wouldn't be surprised to see it cut the CPUs yet again by half or thirds. Depends on how big & complicated the service in question is.
I guess you could even use that as a metric... if someone come up to you and said "I've got a magic button that if I push it will cut your code's CPU usage in half. How much will you pay me to push it?" and if the answer is a non-committal shrug, Go's a fine choice. I have about a dozen Go services and I'd pay you about a buck to push that button, because they're already way more efficient than I need. Uber would clearly pay quite a bit.
But mostly it depends on whether the pacer would perform a minor or a full collection in that scheme./
Now, granted, 2021 is probably down on that further but it's still a lot.
I believe Uber is also well known for building everything in house as opposed to using common cloud services. So they need to run their stuff.
By the time they migrated they already were mature in the market.
Is there any advantage using a 'time.Time'? This is just a simplification to make the example less obscure? Or the change is so insignificant(%-wise) that it makes no sense to optimize for it?
There is tons of great gc research/experiments/learnings out there that can have real benefits, maybe the JVM with its zillions of knobs is a step too far but find a middle ground seems like it would help people.
That's not to say tuning isn't valuable or necessary, but that the vast, vast majority of Go programs will never need tuning. I cannot in good faith say the same about JVM programs, even with the most recent and modern GC profiles.
Other than possibly max heap size, G1 should only ever be tuned by the target pause time value, which chooses between latency and throughput.
Boxed/unboxed I agree could take a bit of love but that is happening with Valhalla.
Yes and that is due to boxing/unboxing. This is also being worked on.
And while you indeed can’t check the generic type of a generic object, I really rarely see any reason for that. Like, if you have written code in any language without reflection, you can’t do that for any object and it is not a hindrance in itself.