The case of a leaky goroutine
brainbaking.com
brainbaking.com
An app I work on recently had a bug where goroutines would slowly build up over time. Turns out the bug is in the Growthbook SDK [1]. We can monitor the number of goroutines, but having a large number of goroutines waiting in the location that gets stuck is normal — we can only see such a problem over multiple days, in that the minimum value slowly goes up.
If Go could tell you the timestamp of the oldest goroutines as part of the pprof dump, we could have an alert, and it would work for any such leak.
Edit: Also, maybe the tool at this comment could've helped you? https://news.ycombinator.com/item?id=39817775
I don't think Goleak would have helped here, because I believe it doesn't support concurrency. It's really designed to run in tests, not in production. It parses and searches stack traces, so it's not going to be performant.
A thread leak can lock up your entire node, including all the control plane processes. A container spec doesn't provide an easy way to control thread/nproc/ulimit limits (you can still do it, but it's not straightforward), which in turns leaves pretty much every k8s deployment misconfigured and vulnerable to thread leaks.
Anyway, we’re using golang for some stuff at work, and holy crap, I forgot how terrible it was to work in high level languages that don’t statically check for correct synchronization.
If C++-style concurrency is like a chainsaw, then golang concurrency is like a chainsaw in a bouncy castle.
> programmers are strongly encouraged to use appropriate synchronization to avoid data races
Any time you need to "encourage" programmers to do the right thing, you have already failed in your language design.
And I think OP agrees with me here. OP says "static checking of correct synchronization" which is irresponsibly absent from Go.
While it does not have parallelism, the concurrency is as unrestricted as it is in go, forgetting awaits is a very common issue leading to wild tasks running unbounded and unchecked, and when you need any other synchronisation mechanism you have to write them yourself and they’re tricky indeed.
$dayjob’s JS team keeps chasing concurrency issues, and breaking builds because of them.
Go does neither. That's why there's this thread on HN that I bookmarked: https://news.ycombinator.com/item?id=31698503
Single threaded languages are usually data race free. I imagine some multithreaded, purely functional languages are too (everything is immutable, and therefore cannot be modified in race with reads). Of course, SQL running in strictly serializable mode is too.
Of those, the only one of those that's an appropriate choice for systems software development is Rust. The Core C++ Guidelines are a runner up in my opinion: They dictate a subset of C++ that is safer, with the goal of backporting the Rust memory safety properties to C++. Swift has also done a lot in this space.
There’s no other way given the leakpocalypse decision. You’d need an entirely new leak-proof language to fix that, and that means you need alternatives for Rc and Arc (or a way to prevent them creating a cycle).
All the code I write is async, so the borrow checker is effectively broken for me. (Wrapping everything in Arc creates weird false sharing at runtime, and I don't want to spend time debugging that class of performance nonsense.)
Do you mean futures that aren't polled to completion, tasks that aren't joined, or literal memory leaks that happen to own futures?
You can start polling then do std::mem::forget on the future. At that point, the borrow checker thinks the future no longer exists. So, it is unsound to pass a reference with a bounded lifetime into a future (which is why you need all the references to be 'static if you pass something into a spawn, or you need to spray Arc everywhere).
It is not unsound for a future to own a reference (in fact it's super common - how else would asnc methods work?). If you leak the future, it can't be polled and it can't be dropped, so any references will never be dereferenced. But also that's a pretty contrived example.
Like you could call std::mem::forget on a future that owns a tokio::sync::MutexGuard and then you'd have some problems with deadlock... but that's not an async issue, it's always incorrect to leak an RAII guard (same as if you leaked std::sync::MutexGuard)
tokio::spawn has a 'static bound because it simplifies things, not because there's some fundamental limitation of futures owning references.
I was liberal in usage with goroutines and channels. But after that experience, I decided to religiously track each goroutine in my head. If I reached the point of not being able to mentally map all goroutines running, I would cut out go. I started using callbacks more too to avoid the pernicious blocking I have experienced.
> PID limiting is a an important sibling to compute resource requests and limits. However, you specify it in a different way: rather than defining a Pod's resource limit in the .spec for a Pod, you configure the limit as a setting on the kubelet. Pod-defined PID limits are not currently supported.
Not necessarilly true if you're using cgo
As I understand it, a cgo function call yields the caller goroutine and then runs the C code directly on the underlying thread, blocking it from use by goroutines. When the function returns, the thread is freed up, and goroutines including the caller can be scheduled again. I'm not sure if the caller is guaranteed to run next, or if other goroutines can crowd it out. I would imagine it probably does run first, if only to receive the return values from the C function and then release its thread affinity. This whole process is notable for introducing some overhead to cgo calls which can be significant if cgo is used frequently.
While you can create new threads in C and thus create thread leaks that way, I don't think any of those threads will be used by the goroutine scheduler, which sticks with the pool of threads it manages.
EDIT: reading the runtime docs, it seems that GOMAXPROCS is not as hard of a limit as I thought:
> The GOMAXPROCS variable limits the number of operating system threads that can execute user-level Go code simultaneously. There is no limit to the number of threads that can be blocked in system calls on behalf of Go code; those do not count against the GOMAXPROCS limit. This package's GOMAXPROCS function queries and changes the limit.
I think cgo calls count as "system calls on behalf of Go code" for this purpose. Thus if you have GOMAXPROCS=1 and more than one goroutine and you make a cgo call from one of them, the scheduler may create a new thread so the other goroutines can still run. You don't need cgo to do this though, syscalls (explicitly or through Go's stdlib) can exhibit the same behavior.
So I think it is possible to leak threads this way, but to do so you would need to spawn goroutines calling cgo faster than the C code can return.
The entire `async` package is built on this. The `race` combinator is an especially cool application.
After doing a big project in Golang, I appreciated this more. We had our fair share of goroutine leaks.
[1] iirc, if the thread is not blocking on a syscall or allocating memory, it will not be yielding to the RTS.
It feels a lot like the monadic composition of Haskell even if the means it achieves it are very different.
I've seen the light of async Rust and I believe (heh, sorry for the semi-flippant sarcastic remark here but I do genuinely love it when I do things right with async Rust) but I feel that writing bad async code is also not something that the compiler will actively dissuade you from. It goes to certain lengths but not quite far enough IMO.
I didn't want to do it but at one point I started reading how is async implemented and that actually lifted a big part of the mystical veil and helped me understand it better. Now if I can also completely internalize the lifetime semantics combined with async I'd be very proud of myself. (But it doesn't help that I am not working with Rust currently.)
Ah yeah it's here.
You don't have to use it, but when you do, it helps ensure that you organize your code in a way that becomes very easily testable. It enforces a modular approach to composing together golang services.
Being more easily testable helps prevent bugs, like these leaky goroutines.
I'm curious though, when do you reach for `fx` in a non-industrial project and when do you not? I still use the same patterns of separating out the implementation from the interface but I've been wiring in the dependencies by hand. I'm curious if folks reach for `fx` immediately or if it's something that requires thought to add. There's also Google's wire library [1] that does similar stuff but takes a compile time approach so it's a little easier to reason about if struct initialization screws up due to weird implicit things.
I still wire dependencies up by hand, but I'm curious what others do.
I realized quickly it wasn't absolutely necessary to use it since most people just make a package in golang, in order to get the separation they need. But when I started using it more and more, I noticed that taking advantage of the DI features of fx, also ensured that I wrote code that had clear separation of concerns.
In golang, it is too easy to just cross package include `new` things you need right in the function, instead of passing it in as an argument. This of course, makes it much harder to write tests for since you can't mock what you need easily.
The binary I built was distributed across tens of thousands of servers in multiple data centers, and had to run perfectly on every release as it took a lot of time/effort to even do updates. This meant comprehensive testing before deployment, so I wanted to optimize my unit/integration tests as much as possible.
I'm not sure it would be necessary for just simple api endpoint microservices, but for a complicated application binary that needs perfect testing, I can't imagine writing golang code without it. The benefits far outweigh the negatives.
Maybe I just don’t know what I’m missing.
Go designers specifically avoided this model (of having a goroutine "id") for reasons I don't remember any more (may be to avoid making them heavier weight?) but this would be one way to stop leaky goroutines.
I think Go has the internal plumbing to theoretically support this, though it might require inserting checks more often. Another way would be to make contexts first-class and automatically insert context checks even when not done (e.g. selects). And also all I/O has to be cancellable.
I suspect Go's designers prefers the current way in which cancellation is explicit.
Agreed. The idea is to panic() if a goroutine has to be forcibly terminated due to GC, instead of a slow leak. Requires more thought though.
Or even one ctx per goroutine and cancel them dynamically according to whatever logic.
It just requires programmer cooperation, but as long as you pass ctx all through the stack down and handle err on the way back, it is not often you deal with it explicitly.
It feels like something similar should be baked in, and inherited from its parent by default(but overrideable). And a cancel would cancel the callstack. Would be nice to make this cleaner in go 2, I think.
Background: I am coming from the JS/TS/Node world, and have decided to jump onto a compiled language. I narrowed down my choices to Go and Rust and eventually decided to go with Rust, because it didn't use GC for memory management.
In fact, since it doesn't have a GC, you can trivially create a memory leak by creating a reference loop... though the ownership checker makes that in itself really difficult, and so it's again less likely to happen than it otherwise would be. At the cost of loops being hard to make even if you want them.
It is possible to do it by using Arc / Rc and cyclical references (which the borrow checker makes infuriatingly difficult -- for good reason!) but you really have to go out of your way for it. And while there are real projects where you need such idioms I have found that you should not try to twist Rust's arm and just opt for something like arena allocator.
> The defer close() seems to close well, but it’s on the wrong channel.
The `done` input channel is supposed to be closed by a caller, and the goroutine is closing the output channel, surely that's the point?
Now from what I know of go channels and understand of the code involved, the `done` channel may never get closed by the parents (and can never get closed at all if it's nil?), in which case the goroutine never receives a signal, never terminates, and leaks. But the explanation below the snippet confuses me completely.
And if that's that... what's the fix? Aside from not doing this sort of conversions? Just "git good scrub", try to make sure you don't rely on cancellation for progress, and hope you don't use raw background channels again, and don't forget to cancel your non-background channels?
This is not a coroutine at all, calling them goroutines was a clever hacker pun along the lines of "GNU's Not Unix". If you treat a preëmptively-scheduled primitive as though it's a cooperatively-scheduled primitive, you're going to have a bad time.
Goroutines are threads, basically, with all the memory-management headaches that implies. Caveat emptor.
This culminated in 1.14, which made the runtime preemptive on most platforms (https://go.dev/doc/go1.14#runtime) in order to fix the last sticking point where a goroutine might not yield: a tight numerical loop might never yield.
This was an issue, because the GC relied on scheduling to slip in its STW pauses, so the GC would trigger STW, progressively pause every goroutine reaching a yield point, but would be unable to ever pause the last goroutine, and the program would pretty much grind to a halt until it was done.
There are ways to handle this (e.g. insert trapping reads in various control structures), but ultimately preemption was considered a better and more useful solution.
¹ Same thing in frameworks like RxJS, where observers in some sense take on the role of coroutines.
There's a library for it: https://github.com/sourcegraph/conc
But this goes to one of the things I've been kind of banging on about languages, which is that if it's not in the language, or at least the standard library right at the beginning, sometimes it almost might as well not exist. Sometimes a new language can be valuable, even if it has no "new" language features, just to get a chance to reboot the standard library it has and push for patterns that older languages are theoretically capable of, but they just don't play well with any of the libraries in the language. Having it as a much-later 3rd party library just isn't good enough.
(In fact if I ever saw a new language start up and that was basically its pitch, I'd be very intrigued; it would show a lot of maturity in the language designer.)
TaskGroup in the standard library - https://developer.apple.com/documentation/swift/taskgroup
Explore structured concurrency in Swift (WWDC 21) - https://developer.apple.com/videos/play/wwdc2021/10134/
It's fairly new; the thing (and I think you address it too) is that the pattern did not exist yet when Go was introduced. Go is averse to adding more things to its standard library, or indeed changing its core fundamentals; I think it's better to have one well-defined way of doing things in a language, instead of adding the mental overhead of deciding between one or the other.
And I doubt Go will remove support for their original concurrency, like, ever. I'd love to see forks of Go made with core elements (like concurrency) swapped out though.
That specific formulation is new but the concept had been floating around for a while. For instance Rust originally got scoped thread in 2015 (before they had to be removed for being unsound).
Part of the reason I rhapsodize about new languages off of that observation is precisely that Go can't add it. It almost wouldn't even matter if they tried to put it in the language proper, because by backwards compatibility the old ways would still work, and it would take a very long time to get the entire ecosystem to the new way.
That was the hypothesis of Golang, for sure.
I think we've seen that it is true in specific contexts. It seems like it's been particularly valuable to massive teams with less-experienced developers coming from academic computer science backgrounds, frequent turnover and high-coordination projects.
I don't think that property has proven to be valuable more globally. Consistency for consistency's sake is particularly costly when the "consistent" solution has significant downsides in some contexts.
On small teams, having consistency result from actual alignment is incredibly valuable, and a sign of a high-performing team. In those contexts I haven't seen consistency itself, enforced by an outside group and especially without a way to work around it when the "consistent" approach has a reason it sucks for a particular application, be similarly valuable.
I think we've seen two ways languages have successfully been open to evolution:
Java was specifically designed to allow it to build on the standard library over time in fully backwards-compatible ways. It has required the central governance committee to adopt proposals, because reflection is slow and poorly-optimized, but it is a far more fully-featured language today than it was an inception without ever needing a reset. By keeping the surface area small & the strong "Everything Is An Object" paradigm in place, it has had remarkable longevity and has avoided the Python versioning pain.
The second are the "sharp knives" languages: Javascript, Ruby and to a lesser extent C++ (only because DLL hell is very real).
All three of these can be used to write software in any paradigm (including Aspects, if one is a masochist), and so let engineers invest in their own productivity. Languages where the standard libraries are indistinguishable from custom libraries require more skill and collective team alignment to use productively and safely, but also allow for solutions highly-opinionated languages can't support.
I.e. unstructured concurrency is like GOTOs. Not necessarily wrong, but certainly nerve-wracking.
I prefer multi threaded programming. Everything is off to the races (pun intended), you need to think long and hard about lifecycle management, who creates a resource, who will clean it up, how do you synchronize, where do you synchronize. It might be hard to get right sometimes. But the concept is simple. The tools you have available are simple. There's no "well everything works fine automagically unless it doesn't because these 10 lines of code".
From the other comments here it’s not clear that there is a modern language that doesn’t have this problem.
By hiding the complexity of multi-threaded programming behind these facades, we let people write multi-threaded programs without understanding what it is doing. That in turn leads people to believe they don't need to understand parallelism to write thread-safe programs well, and it simply isn't true.
Companies don't want to pay for actual expertise, so we are all pressured to do things we don't understand. Languages can enable us to succeed some of the time, but they can't make distributed systems actually simple. The errors that pop up as a result are a predictable consequences of relying on under-trained workers without the resources we need to do our jobs well.
Then you had sheltered and privileged career which I am finding out applies to a lot of folks on HN, apparently. So hold on to your cushy job because if you leave it you might find out the world has moved on long ago.
Especially Rust's async/await are instrumental to services that literally saturate their network link and would go higher if the link allowed more bandwidth. Normal multi-threaded code is folding under pressure somewhere at the 50_000 requests / sec mark on an average small-ish VPS. In the meantime the async/await code is saturating a 10GbE link.
Both Golang and Rust didn't make the perfect design choices though -- that much is true, sadly. But they are a big improvement over the status quo.
I'll grant you that Golang replaced one class of problems with another -- something I dislike as well. Wish it was stricter but it's good enough for 95% of all projects everywhere, if we have to be brutally honest with ourselves.
There are a small number of very large companies where their architecture is based on that sort of thing. But in most cases in most places the right approach in that situation is to ask, "why is this a thing you think you need?" And then to either inline the service or employ horizontal scaling strategies instead of trying to stuff more cycles onto one network card.
I do agree that these language choices are very likely a symptom of the misuse of service architectures. People are treating services like objects (or worse, Singletons) that live on a different machine, relying on the network stack as their interpreter. It makes "modern" software buggy and fragile. Neither of these languages solve that problem, but they do capitalize on it.
Very strange hill to die on, and feels like a side attack towards a point that you chose to ignore: namely that classic multi-threading is fragile and does not scale well. Doing the async/await thing is objectively better. I get why people don't want to admit that they invested so much in something and are grumpy that this skill is now not as needed -- I was one of them in fact, I am coming from C pthreads and Java after -- but being a curmudgeon and trying to do a torpedo attack on a discussion about multi-threading vs. async/await is something that I can't endorse.
> "why is this a thing you think you need?" And then to either inline the service or employ horizontal scaling strategies instead of trying to stuff more cycles onto one network card.
Very often false, having stuff being only on one machine with 1-2 hot backups and a load balancer is the lowest maintenance I've done on systems for all of my 22 years of career. Feel free to disagree but (1) most programmers will never work on the scale of AWS and Facebook and (2) horizontal scaling / distribution comes with many, many new failure modes.
KISS is an art that is being forgotten, it seems.
async/await is particularly damaging because it breaks the paradigm of Javascript. That no one can draw a picture explaining it is a huge problem, and it encourages people to write callback hell code by hiding how ugly it is. Unfortunately, whether the code is pretty or ugly the massive webs of nested callbacks are still a source of massive complexity and potential failures.
I constantly see tests in the wild that are passing even though the assertions fail because people don't understand the concurrency model they are using. And at least 60% of the time there was no reason for that concurrency to exist in the first place, except that it's what the code example looks like.
To be fair, I do look forward to when logic programming languages get their time in The Sun.
But if we can go to (X/2)% error rate then that's still a win.
I wouldn't mind if we replace Golang and Rust in 10-ish years or so. For now they are definitely doing better than C++, especially having in mind that the old guard is gradually retiring and the newer generation are not as good with it.
You seem disappointed that we haven't found the one true universal language yet. I am as well, but no need to trash-talk the current iterative improvements. Apparently that's how we'll get to that ultimate thing.