Practices for writing high-performance Go
github.com
github.com
I have a question that I don't remember being answered in any tutorial I've done so far. I've written a lot of C code and I typically make memory managed lists, so if I need a new common object I grab one from the list to avoid free/malloc as much as possible. Does Go do this automatically, or should I still do this on my own? I'm writing a long-running server, not a utility or something short lived.
As of now[2], they get cleared when a GC occurs so they can be a bit quirky to use in practice. This behavior is changing though
[0] https://golang.org/pkg/sync/#Pool
I don't know if the Go runtime has optimizations for this kind of thing inherently, and it might.
Profile, and see.
Go's optimization for this is escape analysis which is a best-effort attempt to put things on the stack. Because the stack analysis is naive (which isn't necessarily a bad thing), it means lots of things are still allocated on the heap, and those are the allocations I'm referring to.
Look at the -m gcflag to see what is escaping, then see if that is where the allocation hotspot is. If so see if you can get the escape to happen where you want.
Otherwise you need to go no allocation. Notice that go actually gets _worse_ over time generally at heap allocation so it’s likrly worse than early profiles suggest.
I don't think it's especially hard to reason about in the main cases, and when you're unsure (as you mention) you can actually print out the escapes. Would be really interesting to have an editor that would show you where things were escaping (although that could easily lead to premature optimization).
I work in C++. I'm sometimes surprised at the way C++ treats its users. There seems to be a culture of pre-optimization. It's hard to write the most abstract application code without constantly thinking about performance under the surface. That's baked right in. Concepts which are used everyday by programmers can barely be cleanly summarized in a paragraph, even by well respected luminaries of the field. Trivia you "just have to remember," impinges on almost every single function or method you write.
That said, there's also a lot of awesome things in C++. It just represents a particular set of cost/benefit dials. Golang represents another.
It breaks its own rule occasionally -- RTTI, exceptions, [0] standard library machinery with always-on thread safety -- but the C++ folks go to extreme lengths in the name of performance.
There's no small irony in the way embedded folks write off C++ as too heavyweight, given that C++ tortures itself in the name of not forcing bloat upon the programmer.
[0] https://llvm.org/docs/CodingStandards.html#do-not-use-rtti-o...
"You don't pay for it," should really be, you == the-CPU doesn't pay for it. On the other hand you == the-programmer has to think about it all the time. For some domains/contexts, yes this is actually very desirable. In that case C++ becomes a performance/optimization Swiss army knife.
Other people might want it dialed in for, "You don't have to think too much about it, until it's time to optimize, and then you can't optimize all the way, but you get enough of the way there."
Visual basic had a Replace() function back in 1998, to replace substrings. C++ now has many amazingly advanced high-level features, but still no built-in way to replace a substring. I don’t think I needed to write my own string replacing function in any other language I used (I’ve been programming for living since 2000).
I like C++ and use it a lot. But these seemingly small issues with its standard library escalate quickly. String handling, IO, localization, date & time, multithreading before C++ 11, many standard collections, and other parts are just not good enough. I pretty much stopped writing complete apps in C++, nowadays I’m using C++ for dll/so components I consume from higher-level languages like C#, Python or golang. And when I do, I often choose to ignore large parts of the standard library in favor of alternatives from atl, eastl, or my own ones.
That's C++ strings, too: https://docs.microsoft.com/en-us/cpp/atl-mfc-shared/referenc...
Not only they have better API (replace, tokenize, implicit cast to const pointers), these strings are often faster. That particular class is Windows-only, but nothing prevented C++ standard folks to come up with conceptually similar cross-platform stuff. BTW, that CString class predates C++ standard library by many years.
> And C++ can quickly become difficult to read after overloading various operators on classes or after using fancier less-known features.
Yes, and I saw quite a lot of code like that.
But in other cases these features help with readability. I often code math-heavy stuff in C++ processing vectors, matrices, quaternions, complex numbers, etc. The ability to implement custom math operators on these structures IMO helps with readability.
What doesn't help is the ability to abuse them, like C++ iostreams do with `operator <<` everywhere.
Not just operators, it's generally too easy to abuse features of the languages, writing code that's very hard to work with. Unfortunately, not doing that requires lots of experience with the language.
I'm not planning to switch due to the good parts. First-party SIMD intrinsics support. Trivial interop with C and C++ libraries: hard requirements like OS kernel APIs and GPU APIs, industry standards like libpng, or just very nice to have like Eigen. Very small runtime allows to build dynamic libraries, consume them from anywhere, and not worry about binary size or runtime dependencies. Also tools like debuggers and profilers are very good.
But when performance is less critical, I'm more productive using other, higher-level languages.
A C dev would start cold sweating has s/he types virtual, or operator something().
And yes, I am. I don't care about earning HN brownie points just to prove something that will be hand waved anyway.
– Rob Pike
The problem they solve looks simple on the outside, the reality is different.
That's probably pretty unusual, but the point is, just because it was written in C++ doesn't mean it's any good. Not everything gets serious attention from people who know what they're doing.
BTW the reason dl.google.com rewrite was faster was not because it was in Go, it was because the C++ server was serving off its local disk and the rewrite was serving off a cluster file system with ~infinite I/O capabilities. Apples and oranges.
This is downvoted grey as I write this, and as someone who has been called a "Go shill" on occasion... it's exactly right. Go is a decent language for writing code in a fairly straightforward manner and getting pretty good performance out of it, but if you need to squeeze every bit of performance out of your hardware, it's a bad choice. (I can give you choices that are worse by an order of magnitude, or even more in some cases, but it's still a bad choice.) You have a fairly smooth optimization ride up to 1.5-2x slower than C for most use cases, a few pathological edge cases where it's grossly worse (many clustered around these "every drop of performance" problems!) and a few where it'll reach parity, and then you're going to hit a brick wall.
An 88 core system is probably not impossible to sensibly use with Go, but you're going to be more constrained. I'd imagine it can probably serve web requests really well, but it's more likely to hit pathological cases if you hammer certain global resources.
Arguably, precisely part of the point of Go was that C++ makes you pay for that level of performance in code complexity and cognitive overhead all the time, even when you don't remotely need it. (I can't prove this, but I'd guess the median "cloud service" is grotesquely overprovisioned on the smallest AWS instance. The "cloud services" that leap to mind are things like the AWS auth servers or Netflix content servers or the Google crawling or indexing servers, but while those are huge and important, they're also in many dimensions the exceptions. A good chunk of the popularity of "serverless" is probably a result of this.) When you do need every bit of performance, though, the list of viable options is short.
I think the point about concurrency typifies this, channels aren't as fast as mutexs and semaphores, but it makes sharing code with co workers easier to reason about.
If you're in a domain where performance is still king, Go isn't trying to find a place there.
regexredux program is outlier in Rust, because replacement of a regex in string is slower in regex crate, because author of regex crate chose to implement safer, but slower algorithm. To fix this, regex crate must be updated or replaced. I spent two weekends on this.
[1]: https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
I somehow doubt pcre2 is being changed to make the tiny toy C programs run better.
> I spent two weekends on this.
So shouldn't we assume the program performance simply reflects all-those-hours you've spent working on it?
fn find_replaced_sequence_length(sequence: String) -> usize {
// Replace the following patterns, one at a time:
let substs = vec![
("tHa[Nt]", "<4>"),
("aND|caN|Ha[DS]|WaS", "<3>"),
("a[NSt]|BY", "<2>"),
("<[^>]*>", "|"),
("\\|[^|][^|]*\\|", "-"),
];
// Perform the replacements in the sequence:
substs
.iter()
.fold(sequence, |s, (re, replacement)| {
regex(re)
.replace_all(&s, NoExpand(replacement)).into_owned()
}).len()
}
It measures performance of RE engine. I can switch from regex crate to PCRE2, and program performance will match C.Perhaps it would; those measurements have not been made.
What does that have to do with re-writing libraries to make tiny toy programs run better?
What does that have to do with program performance being a proxy for programmer effort?
Also, there might be a something to prove "bias" :-)
Graph analysis and/or routing, for one. Distributed Dijkstra is pretty impractical (the optimal lower bound on the number of messages required equals the number of edges in the graph!)
For example, I don't work for Google, but I can pretty much guarantee you that an individual Google Maps routing query runs on a single, extremely-high-memory (but not NUMA!) node, which has likely been optimized for speed using plain-old threading on a high-core-count CPU, with the upper bound on the number of useful cores being the system architecture's per-socket memory bandwidth.
https://www.nextplatform.com/2017/06/22/casing-hpc-market-ha...
Although they cost a ton, there were at least two reasons to like NUMA machines:
1. You program them much like multithreaded machines instead of message passing like MPI. You do need to account for locality. There's OS, library, and documentation support for handling that, though. Porting a multithreaded library to NUMA is much smaller problem than clustering it.
2. A massive amount of memory with lower-latency access than on clusters. If your data is in memory, it's much faster than if it's being moved in and out of memory. Then, when it is moved, it's moved faster. Low-latency reduces the damage of many smaller copies, too.
For such reasons, I always wanted a SGI or Cray machine. Modern, multi-core machines with plenty of RAM are good enough for most of my purposes, though. I do plan to do some model-checking of software in the near future. The amount of RAM used grows exponentially or something like that with the size of the program. Small ones already use GB of RAM in the analyses. Some are getting parallelized a bit. Obviously, 88 cores with 100+GB of RAM could be pretty useful if handling programs twice or three times as large. :)
Btw, there are also languages designed specifically to take advantage of parallelism in many HPC situations. Like Go, they were intended to let you describe the algorithms in a high-level way with the compiler synthesizing efficient implementations for everything from multi-cores to NUMA to clusters. That's the theory. The best one from my prior research was Chapel. Then, there's simpler ones for stuff like data parallel with Cilk being an example.
https://en.wikipedia.org/wiki/Cilk
ParaSail was a recent one with interesting design. I'm not sure what its current status is in terms of usability.
https://www.embedded.com/design/programming-languages-and-to...
If you want to saturate every core on a very parallel problem like "handle as many packets as you can", a very rough first approach would be to spin up 88 threads (one per core). Within each thread, you could use either something like boost::Fiber[2] (similar to goroutines, except mapped N:1 to OS threads, rather than M:N) to avoid blocking the threads on IO. This paper[3] has a good overview of different concurrency approaches.
If you're doing something that's not embarrassingly parallel[4], then often there is a well-researched approach for your specific domain. The same general ideas apply to keeping your cores busy, but you're often more bounded by communication between threads, memory bandwidth, etc.
[1] https://en.wikipedia.org/wiki/POSIX_Threads
[2] https://www.boost.org/doc/libs/1_70_0/libs/fiber/doc/html/fi...
Profile your channel latency, then design around that. "Small RPCs?" With an 88 core server, should you be making that many remote procedure calls? Find the 20% most intensive code and take that out of the hands of the scheduler, as much as possible.
and runtime.mallocgc
Even with sync.Pool? Can you design those parts of the system to mostly use the stack?
You can very carefully design around it using only channels but it quickly starts making more sense to just use a different concurrency abstraction.
They aren't as efficient, but using them as a semaphore (i.e. only signaling) can facilitate fairly-efficient, easier-to-read shared memory. But I'm kind of proving the whole...
> You can very carefully design around it only channels
...part of your post.
It gives you a simplified fork in goroutines, a way to safely pass by value or signal between these through channels, and structures like select to psuedo randomly deal with race conditions.
That stuff is all awesome!
But when you are talking about performance, opinionated simplicity is a bad philosophy. You are going to want the finer grained control the poster is asking for, and which Go deliberately doesn't allow for. C, C++, and Rust will always have a speed advantage because they lack the overhead of the GC and were designed with allowing the developer minute control of the system, which is the space to make such optimizations.
That said, opininated simplicity means every Go codebase I drop into looks roughly the same. That is not true in the above langauges, precisely because they allow for more control.
To me, Go's ideal use case is the domains of Java, Python, NodeJS, and even Elixir -- where it has a performance edge in many cases and where its consistant style really shine.
All that to say, Go is bad at performance optimized concurrency, as designed.
https://github.com/golang/go/issues
It should probably explain why one cannot "avoid its global allocator" by using pools and/or stack objects.
sync.Pool is fun but it has known scalability limits that are yet to be addressed in a released runtime (see https://go-review.googlesource.com/c/go/+/166960/ for example). You also can't use sync.Pool for anything that you need to be non-ephemeral, like CPU-local counters, because the GC can just blow them away at any time and the finalizers run at arbitrary future times or possibly never.
All of your listed issues have additional details and discussions around them. It not "we don't add this stuff because you are too dumb to use it". It's more "we don't know how to add this currently so it would properly interact with other features and be non-confusing". There are some exceptions to this rule (time.Time issue comes to mind) but overall they tend to listen to the developer base. Also, keep in mind that there are also relatively few Core Go developers, so they have to be extremely careful about what they gonna add in the language. It's going to be them who is going to support it in years to come.
As for C++ - yes, nothing beats it in highly optimized cases. And there is nothing wrong with that. The problem is not only writing C++ for those highly optimized cases is an extremely hard task by itself. The problem is finding people who can actually do it with the resulting code which is going to be supportable in the future (I'm not talking about UB, races or memory leaks).
[0]: https://github.com/golang/go/issues/18138
Edit: wording
I was hoping for more insights about Go specifically.
And I think that's the point. The language doesn't get in your way, and it's not any specific language feature that will make your code go fast or slow - unless you abuse it, or use it when you don't need to.
I mean if you're happy with C then by all means; Go isn't claiming to be faster than other languages, not this close to the metal. It's mainly aimed to take away some of the mental overhead you get with C and similar languages. And compile times.
in practice I haven’t seen any major performance overhead when passing values about.
Almost all types that could hold a lot of data have reference semantics, so if you have a struct with a large slice, or a large map, all of the actual data is going to be on the heap, and not affect copy performance.
The only time you might want to consider using a pointer to improve copy performance is if you have lots of non-heap-allocated data in your struct, and in practice the only way that happens is via large arrays. That is, e.g. `[8192]string`, and not simply `[]string` (which is an efficient slice). But you should never do it preemptively, you should always benchmark both styles and only switch to pointers once you have proof it's meaningfully affecting performance.
Not that you should not have mutexes in idiomatic Go, but they usually have to be warranted for.