Go memory ballast: How I learnt to stop worrying and love the heap
blog.twitch.tv
blog.twitch.tv
Until Twitch hit the bottleneck with their service, they were reaping multi-year benefits of faster development using memory-safe language (huge gain for security).
Their short- (or mid-) term solution seems decent and more importantly Go core team is working on addressing this particular issue (which given the Go 1.x backward compatibility promise makes it really easy to reap all the benefits with each Go release).
Tiny nitpick. On Linux at least you can read from the slice. New memory allocations are given a copy on write zero page (full of zeroes). So you can easily read that 10GB worth of zeroes out and still memory consumption wouldn't increase.
Only when you write a pagefault is issued and you'll get real backing for it.
Agree the JVM is complicated, but you don’t have to be that complicated. A lot of it is legacy anyway.
The way a gaming engine would do it would be to disable the GC and allocate a large chunk of memory and read/write directly to that memory.
Its a pretty quick hack, not bad, really.
I'm also happy to make my code less allocation-happy or more GC-friendly in the extremely rare circumstances when it matters. Like this article!
But I think we probably just work in different environments.
It was a deliberate choice not to have as many options as Java.
Someone should let Microsoft know they need to rename their “C++ Runtime Library” package.
However, manual memory management makes development more expensive. Especially if you don't crunch numbers but parse strings. C++ probably ain't a good replacement for Go. For the last decades people and whole industries have been migrating the other way, from C++ to higher level memory safe languages, initially Java and C# then others followed.
I avoided mentioning that the C++ runtime was optional because I was trying not to be pedantic and I assumed in context it would be understood that this was irrelevant. I think I should’ve just made my point earlier, but:
- This all started with someone mentioning that C++ doesn’t have ‘bad runtime behavior’ because it doesn’t have a runtime, but this is false. It has a runtime, or to be exact, a specification for one, and I think it is safe to say a vast, perhaps extreme, majority of C++ developers are using it.
- If you opt to not use the runtime, then your runtime behavior is dictated somewhere else, but runtime behaviors don’t go away. In freestanding, you may be your own runtime, but then your runtime behavior is defined by the machine you’re running on. You just get to choose what layer of abstraction you are sitting on.
- edit: And also, as a point I forgot to mention initially, I don’t really feel like freestanding C++ vs standard Go is an apples-to-apples comparison.
I think the point was that the C++ has less runtime behavior, since it doesn’t have a scheduler or garbage collector, but extrapolating that to no runtime is wrong even if the runtime is optional.
BTW, C++ does have a scheduler. Optional like the rest of the runtime, but it’s 1 line of code away in all modern compilers, that line starts with #pragma omp.
However, even if you go to the extreme of not using any features that require runtime support, if you are using hosted C++, in practice you still have one bit of runtime: the entrypoint. Technically, an operating system could implement the bits that call main, but to the best of my knowledge none of them ever have. So every compiled binary from every hosted C++ implementation begins at the runtime library. (Admittedly, normally the C runtime library, since C++ doesn’t differ here, but that is just another layer deep of pedantics. In practice, everyone has a runtime.)
Nitpick: As far as I am aware OpenMP is not part of the language itself but an extension. But yeah, you could argue that is runtime scheduling in C++, I think.
The really obnoxious ones are std::stoi, stol, etc. which were added in C++11 as a replacement for strtol and friends, but throw on overflow/invalid input instead of returning a result type or a bool with out parameter.
And they don't even have a workaround for the GC assist being way too aggressive, they're just sucking it up and letting the latency exist.
No
Weak argument against Go ?
No
Let me know which language was perfect
That there is a proposal to work this out and relevant people are working on this makes a strong argument for go IMO.
But further, it seems fairly easy to fix and indeed it looks like the Go team are doing just that.
It references a patch to add a SetMaxHeap call, which could help in these kinds of situations too.
A blog post from someone working on .NET suggested "the user shouldn't have to know about GC internals" as the heuristic for when you have too many knobs, which I like. The user knows some things that the runtime does not about their needs in terms of RAM vs. CPU use, and it seems reasonable to have ways to communicate that.
(One other thing the user knows is their relative priorities for having a consistent level of CPU used by the GC vs. minimizing the absolute amount of CPU used. I think Go is hard-coded to target about 25% CPU use during a collection right now. Not sure we're ever getting that knob, though.)
How did they allocate that 10GB byte array?
func main() {
// Create a large heap allocation of 10 GiB
ballast := make([]byte, 10<<30)
// Application execution continues
}It seems to me that if you're going to play with memory directly, it's reasonable for the generally-memory-safe runtime to throw out its guarantees on memory-safety
There are ways that could be done (for example by making the type system aware of whether or not something is hand-managed and potentially unsafe) but not without adding considerable complexity to the language and, presumably, the implementation.
This commonly recurring idea that we need to somehow banish all insecurity and risk is flawed in my opinion. When safe is the easiest and most convenient way of doing something, it will generally be used. What we need are sane defaults, combined with explicit and unambiguous interfaces for when we do choose to take manual control. The real issues start to arise when a language exhibits unexpected behavior, when automatic memory management is opt-in instead of opt-out, and similar.
You can do something similar for go where you have a function that marks your locally-scoped variable dead, the compiler can track to make sure you don't use it afterwards, and if it's not the last reference to the underlying, it's not deallocated. If it is, it gets culled immediately.
|_|()-Xms sets the initial heap size, but GC happens before the heap is exhausted, in most GC configurations. There is a smaller eden in the heap, and GC is needed to evacuate objects out of the eden into the main heap.
That paragraph got me interested. A lot of stuff in this article is based on this statement. Where can I read more about this? By intuition I would have guessed that the process of marking is the smaller portion of work ...
My guess would be all the pointer chasing during marking is expensive and since Go doesn’t use a compacting GC there isn’t anything to do other than push freed memory back into the allocators internal data structures during a sweep.
Sweeping, in typical implementations, also requires locking/unlocking mutexes and sorting/combining freed memory chunks to combat fragmentation, and that can be slow.
My hunch is that sweeping is fast because there isn't much garbage (meaning distinct allocations, not megabytes) per sweep cycle to start with.
Edit: the talk referenced in the article (https://blog.golang.org/ismmkeynote) provides a hint. Go has lightweight threads, and if you have many of them you have to walk all stacks, closures, and all registers of all threads in order to determine which objects are live. If they are processing millions of requests per second, they might have a goroutine for every request, and that might explain why sweeping is (relatively) expensive in his benchmarks.
What does annoy me is that GC languages doesnt provide option of allocations that are not handled by GC but still uses language primitives and/or privide option to deallocate them manually. Dev. could decide should they leave alocation to GC or handle them on their own. I would certanly love it, in most programs that I write I know exactly what the lifetime of objects is and I am leaving them to the GC just due to missing any other option.
The language has improved itself so much that it's now a better syntax and verbosity
Mesa/Cedar, Modula-2+, Modula-3, Oberon, Oberon-07, Oberon-2, Active Oberon, Zonnon, D, Nim, C#, VB.NET, F#, Standard ML (with MLKit), Eiffel, Swift.
Just the list I tend to keep in mind, there are quite a few others.
Why would you provision a 64GB machine if you only need less than 1% of that?
Especially given earlier:
> One approach to handle this is to keep your fleet permanently over-scaled, but this is wasteful and expensive. To reduce this ever-increasing cost, [..]
This already was over-scaled.
Even if you physically deploy them, that would be a silly loadout. Take the bit of insurance out and stick some RAM in it. The cloud providers don't offer these instances because if you run the numbers they don't really make sense.
If above is true, then I have one question: What if some where down the line, somebody calls append on that buffer?
bigBuffer := make([]byte, 1024)
clientABuffer := bigBuffer[:512]
clientBBuffer := bigBuffer[512:]
clientABuffer = append(clientABuffer, []byte("Hello Client 2")...)
// clientBBuffer will now be [72 101 108 108 111 32 67 108 105 101 110 116 32 50 0 0 .... even we didn't directly modify it.
Could be a downside.No, it's actually just "create a big buffer and don't drop the reference". It's never used for anything except changing some numbers the GC uses to do its logic. Since it doesn't actually end up in physical RAM, it's doesn't consume significant resources either. It's just a funny-looking way at the Go-languange-level to twiddle some numbers to make the GC act differently.
Allocating a big slice and handing out chunks of it does have its uses, basically, arena allocation flavored by being used in a GC'd language. But that's not what this is.
Thank you for sum thing up.
B) In a non-memory managed language, wouldn't you just run into the same problem except it's called "heap fragmentation" instead and malloc has to do an unreasonable amount of work to find free blocks to use?
Go is chock full of "what were they thinking?!" decisions.
That’s not to say it’s all bad, far from it, but these were my wtf moments exploring the language.
This is hardly a weird choice. Garbage collection gives you the highest productivity while being memory safe out of the three options (manual, GC, static).
> lack of proper generics
They are working on parametric polymorphism, since a while now. It is hard to get right in a language which has readability and simplicity as a main focus.
> interface{}
I certainly agree on this one. It is a weird feature and often a 'smell' when found in Go code.
> the weird way of distinguishing visibility via capitalization
It is weird syntax choice in the sense of being unique/uncommon but fits very well into the readability focus of Go.
> comments that affect code generation
I agree that they should have introduced syntax for this, assuming you mean compiler flags (or w/e they're called).
> archaic plan 9 assembler instead of using LLVM
One of the goals of the language is to compile really fast, which they certainly succeed at.
I certainly agree on this one. It is a weird feature and often a 'smell' when found in Go code."
The feature itself is not a problem. Every major static language has the equivalent. It's often called something that involves the word "dynamic".
Having programmed in Go for many years now, I don't find myself using it very often in my own code. I've come to think of this as something said by either people who have never used Go at all, or people who used Go briefly but insisted on programming Javascript-in-Go or something. The latter is definitely a Bad Time... but it's always bad to program X-in-Y. If your code is shot through with interface{}, you either chose Go for something way outside of its domain, or you are not using it correctly.
"> comments that affect code generation
I agree that they should have introduced syntax for this, assuming you mean compiler flags (or w/e they're called)."
This is another criticism that I think mostly comes from people with a checklist criticism set of Go, because in practice, this is of negligible concern. It doesn't come up often, it isn't proliferating (i.e., it's not like with every point release we get another two or three new kinds of comments), it's literally never been an issue of any kind for me in the last six years. It's a complete non-issue. I am far more annoyed by, say, the fact godoc doesn't give me basic markdown than comments affecting compilation has ever annoyed me, and that's just an occasional minor annoyance.
Where it is appropriate, it is a very cool thing that you can have fully dynamically typed variables with all safeties in place. Of course, it shouldn't be used in place of better abstractions, like specific interfaces or properly factored code.
It is an issue as someone who doesn't write Go code often but occasionally reads/debugs it. At least for the first couple of times I've come across it made me scratch my head for a couple of hours, because I didn't see what the issue was it being 'hidden' in the comments. Which is fine. It just isn't in line with the general premise of Go's language design, so I wasn't even expecting it. So in a sense you are right!