Simple techniques to optimise Go programs
stephen.sh
stephen.sh
That said, the example using sync.Pool is not quite right. The New function should always return a pointer type to avoid incurring an additional allocation on Get [0]
The code should look like:
var bufpool = sync.Pool{
New: func() interface{} {
buf := make([]byte, 512)
return &buf
}}
b := *bufpool.Get().(*[]byte)
defer bufpool.Put(&b)
[0] see example in https://golang.org/pkg/sync/#PoolEDIT: Downvoters, I'm terribly curious which of these points you disagree with? Do you believe higher level languages (like Python, JS, or even Go) are faster than lower level languages (like C, C++, etc)? Or do you believe that HLLs are actually faster (or meant to be faster) than LLLs? Or maybe you are downvoting because I don't know what a "higher order language" is?
I think what people are trying to get at in this thread is that they'd be happy if their Go compiles took 0.5 seconds instead of 0.1 seconds, if it also meant that it lessened the need to manually employ the techniques in the OP. Of course, keeping such tradeoffs within reason is a difficult prospect.
I'd like to believe this is the case, but has that ever actually happened - as in a programming language seeing that much speedup in its compiler? There're quite a few cases of languages that are generally taken to have long compile times (C++, Haskell, Scala, etc) but I don't know of a case where any has been able to improve by as much as an order of magnitude.
As for Rust specifically, a benchmark I saw a while back estimated that modern rustc compilation is about 3x faster than it was as of the 1.0 release in 2015. Other "cheating" approaches to reducing build times for Rust are incremental recompilation (which is kind of like built-in ccache), parallel codegen units (the original rustc, ironically, was single-threaded), a compilation mode that stops after typechecking (thereby avoiding codegen and LLVM entirely; this is invoked via `cargo check` and suffices for much of one's interactive development), and an alternative debug-build backend (Cranelift, still very WIP) which is optimized for extreme compilation speed (like a browser JS engine) rather than compilation quality (LLVM being largely designed for the inverse).
Indeed, the first step to making a fast compiler is "design your language from the start such that it can be compiled quickly" (which can be seen in the design of Go). I can think of a few things in Rust that might have been designed slightly differently (in the name of compilation speed) if the authors had the compiler-building wisdom they do today. But there are still plenty of opportunities to improve here.
Also, taking ‘its compiler’ literally: Turbo Pascal was over ten times as fast as earlier pascal compilers for the same hardware, which were multi-pass.
Golang's compiler is fast because they don't check everything they could, the language itself was intentionally dumbed down, and because it doesn't bother dealing with dependencies as tightly (or as well) as other languages. I'd say it is important to know that.
On sync.Pool: good spot. I'll update it - thanks.
I believe the only way to do this without unsafe code is to pool these slice-headers as in https://github.com/libp2p/go-buffer-pool/blob/master/pool.go.
You won't see these aspects of poor design on a sampling profiler. You will see it by running e.g. perf on Linux and seeing pitifully low IPC and cache miss numbers.
This contributes to Rust programs generally having good performance characteristics without spending time on optimizations.
Could you clarify this? It seems like the opposite to me. Borrowing requires lifetimes in type signatures, but boxing yields owned objects, which can be passed around easily like value types.
That being said of course in almost all of these case you can restructure your program so you don't need to box the values but if it's not performance critical why bother? Repeat a couple dozen times across a large codebase and you have the same pointer chasing issues.
Some patterns of writing code will be really awkward to realize, but there are usually "more rusty" solutions that you start to apply without event noticing. Once you write code with the desired ownership semantics in mind, it's often (relatively) frictionless.
Too many people have taken the “make it correct then make it fast” advice too far. The definition of “correct” needs to include performance parameters from the beginning because usually the bottlenecks that cause real issues are architectural.
What I will say is that code that is architecturally correct for the performance requirements necessary is largely not less understandable or unmaintainable than code that is incorrectly designed for the performance context. You don’t end up in the kinds of harder to read optimizations you see in this article in those cases any more than you do under the other case.
Another consideration is if your architecture is wrong for your performance space, nothing will help but a rewrite. If your code is optimized in the small too early, you can always rewrite in a way that is more clear.
Your design should indeed take into account performance requirements. But micro-optimizations like these (almost all of these changes are to avoid linear numbers of reallocations, with the exception of string builder) don't give you order of magnitude speedups unless they're in hot loops anyway.
Profile than optimize means profile, then optimize. Designing good software isn't optimization, its designing good software.
You shouldn't micro-optimize before profiling, because it likely won't matter. Bluntly, if you have a flat profile, none of the optimizations in this article are relevant anyway. You'll be able to pull out single digit percentage speedups, maybe.
The optimize then profile argument isn't meant to be about architecture. Yes, you should build performant architecture. Yes, you should take time to plan a performant architecture before building[+]. But the question of profile than optimize is never (except in the strange way you're bringing it up) about doing macro-optimizations before you've written a line of code. It's almost always in the context of "don't just try to optimize what you think is slow, because you're almost always wrong".
Big-O style speedups from architectural changes aren't micro-optimizations, they generally sit outside of that conversation entirely.
As an aside, flat profiles are in practice exceedingly rare. Most (useful) programs do the same thing many times. Its very unusual to see a program that isn't, in essence, a loop. And the area inside the loop is going to be hot. The pareto principle applies to execution time too.
[+]: Maybe startups who gotta ship it to survive as the exception.
Edited to add: Since apparently you also work at Google, you should walk over to Svilen's desk and just ask him if profiles of production software are generally flat, or if they generally have hot spots.
1: https://static.googleusercontent.com/media/research.google.c...
That's what you get after you profile and optimize.
Glad we agree.
> The optimize then profile argument isn’t meant to be about architecture
Glad we agree. If only all the people who tell me “correct before performant” agreed with us. In practice, in my experience, this is not the case. People use it in day to day conversations at the earliest parts of conversations about architecture all the time. If they didn’t I wouldn’t have nearly the problems I do with the statement.
> As an aside, flat profiles are in practice exceedingly rare
This seems to be the most controversial part of our disagreement. In my experience, that is flatly untrue. Especially when talking about systems where the performance does not meet the requirements. I can count on 1 hand the number of times I’ve seen systems go from “unacceptable” performance to “acceptable” via micro optimizations. I’ve never seen one go to “great”. I don’t know how to quantify this though, so I’m willing to leave this in the realm of my experience is different than yours.
All that is to say, my experience says that systems that don’t treat performance as first class requirements don’t tend to meet their performance expectations.
All of which is neither here nor there based on the article but is directly related to the question of ‘what do you do with a flat profile’?
Also, when profiling production code, it can be hard to FIND the slow function since the optimizer may inline things.
There is perhaps value in understanding your language and ecosystem very well, so that you have decent intuition for what is fast, and making default-fast choices even if their is a small readability cost. That cost may be made up for by not having to crunch on performance later. As well, many performance choices make code simpler, after all the goal is to do less. and less is less.
My comment re write-profile-optimize was geared towards the latter, i.e. using strconv is usually uglier to read than fmt.Sprintf, you probably should leave it alone unless you've profiled and found it to be a problem.
(Note for anyone new to this that the "[]byte-value" - we say "the byte slice" - is a distinct thing from the "values stored-in-the-byte-slice", which is a heap allocated backing array)
While this is good advice, it's not entirely correct. Even with the current go compiler there are ways to use sync.Pool with non-pointer values without incurring in the extra allocation, e.g. using a second sync.Pool to reuse the interface{}. Although I would not recommend it as it's slower, and much less maintainable.
> The safe way to ensure you always zero memory is to do so explicitly:
// reset resets all fields of the AuthenticationResponse before pooling it.
func (a* AuthenticationResponse) reset() {
a.Token = ""
a.UserID = ""
}
I think this is safer, in face of modifications to the AuthenticationResponse structure, and much clearer in its intent: // reset resets all fields of the AuthenticationResponse before pooling it.
func (a* AuthenticationResponse) reset() {
*a = AuthenticationResponse{}
}This would, of course, be much less of an issue with a generational GC, which doesn't have to scan the entire heap on every collection.
Just curious, do vendors on Amazon get reimbursed for the drops in revenue during tests like this?
Should Amazon charge vendors more when they identify where to invest more engineering effort, as a result of these A/B tests, which eventually lead to far higher increased revenue?
Am I the only one who has done this and used append() within a loop, and resulted in a slice that is 2x as long as the original desired length, and the first 1x of items are all empty?
I quickly caught it and fixed the approach (instead of append, I used the `i` as the insertion index into my new empty slice which already had capacity allocated), but the ease with which that subtlety could be overlooked turned me off to this approach unless I profile and really find it's a hot path.
Edit to add: I tried his code, and it resulted in the new slice being as expected.
That will allocate a slice that can fit length ints but has a length of 0 (so append starts at index 0).
You can call len(slice) and cap(slice) to see the difference, but append will insert an element at index len(slice), growing the capacity if necessary
make([]T, 0, n)
and make([]T, n)By default, the slice is filled in and fully initialized. On the one hand, IMHO and IME that's not really what you want if you're going to append things which is a pretty common case, but on the other hand, worrying about "capacity" for slices, while perfectly sensible once you understand the abstraction, is something you can get in trouble with if you're not paying attention, because it makes you really pay attention to the fact the slice has an underlying array.
(Many languages have this concept somewhere, Go just happens to put it in the syntax: A "slice" is an object with a reference to an underlying array, an index of the first element, a length, and a "capacity" which can be smaller than the underlying array has room for. If you want to grow past the underlying array, the runtime then has to create a new one and copy the elements out into that.)
Similarly, if you have a slice of ten elements and you want to take a subset of it and start appending your own bits and pieces, you have to do that carefully: https://play.golang.org/p/QcZOndSp-9x There's a syntax for reducing the capacity of the new slice deliberately, which forces any new "append" to copy to a new slice, producing the behavior you want.
For as tricky as this can get, I'm surprised it doesn't come up more often, but so far in the last five years, this was the source of my bug I believe twice, and one of them I was doing rather inadvisable things anyhow. It turns out that in most "just code" you're not cutting lists up and distributing them various places very often.
Go isn't unique in this; there are many languages with runtimes that require similarly intensive amounts of copying and context conversion before C code can run on whatever the data is. However, most of those languages, like CPython, are themselves slow enough that the penalty isn't as noticeable against the general background noise. (In general CPython requires a lot more copying too; Go is closer to C struct and array semantics and can more often get by with some form of memcpy, the internals of Python look nothing like that.) Go is fast enough that it's much easier to get into scenarios where in a tight loop you're spending 90% on CGo overhead if you're not careful with data flow. For those languages that have to copy a lot out of their runtime and are also fairly fast, they'll face the exact same issues. It's not really a "Go" issue per se, but the challenges faced by any language that wants a runtime significantly different from C.
(One of the miracles of Rust is building an environment and runtime that isn't stuck on C's limitations but at the same time can still speak to C really, really cheaply. Plenty of languages have one or the other of those, but there aren't very many that have both. I'm not sure there's any other language that has threaded that particular needle so cleanly.)
I agree, and some of this overhead is even inherent in legitimate, foundational choices such as VM-interpreted/"managed" code (as in Java/.NET; but Python does this as well) or the use of tracing GC (which requires some strict discipline on heap contents, so as to enable the GC itself to reliably "trace" and discover the semantics it cares about). So, I'm definitely not saying that the choices made as part of Go's design are consistently wrong here!
Indeed, Haskell is in a very similar place overall; the "fibers" that GHC uses in its compiled code are implemented via async code underneath, much like Go's goroutines, and Haskell's performance is also well within an order of magnitude of pure C-like code. So the issues you describe are quite well-understood in that context, and they are definitely not regarded as a "reason" to avoid the use of C FFI when performance requirements call for it. But describing Cgo itself as something that's high-overhead and should not be used for that reason is misleading to an even stronger extent, and that's what I was objecting to in the grandparent comment!
Go is plenty fast for most things, notably things that work well with a GC... but cgo doesn't make the GC go away.