Also, after publishing that blog, they also said they had wanted to try Rust for other reasons:
wanted to try out rust as well for services like this, due to adoption elsewhere in the company. Also, after upgrading between 4 golang versions on this service and noticing it didn't materially change performance, we decided to just spend our time on the rewrite (for fun, and latency) and to get a head start into the asynchronous rust ecosystem. [2]
And nothing wrong with that decision from my point of view (Rust is great!), but I believe they could have solved it without dropping Go for that particular service that they blogged about, at least as far as I was able to understand.
[1] https://blog.discord.com/why-discord-is-switching-from-go-to...
Do you know if that was the case when the decision to switch was made? They said elsewhere they tried Go 1.7-1.10 and none of those solved their issue, and at least back when the blog post was released I couldn't figure out whether the fix for the corresponding issue was known to have been slated for release in the next Go release.
It’s a good question.
The timing is a bit confusing because the blog was apparently published a bit after the fact, but I think I saw that they said they made the decision mid 2019.
Go 1.12 seemed to address their latency issue, which was available as GA in Feb 2019.
That’s based on some retroactive benchmarks on a set of old Go versions starting with 1.9 and targeting what they described as the problem and symptoms.
Things seemed to line up, but I can’t be sure.
I don’t know if Go 1.11 would have addressed their issue.
(In general, Go GC tail latencies including for large heaps have improved a bunch since the last version they said they tried, which I think was Go 1.10).
In any event, they saw a problem, and made a rationale decision for multiple rationale reasons.
> Another Discord engineer chiming in here. I worked on trying to fix these spikes on the Go service for a couple weeks. We did indeed try moving up the latest Go at the time (1.10) but this had no effect.
So I guess that means that there's a reasonable chance that they didn't see a fix on the horizon.
That being said, I have no idea if work to fix this issue was visible on Go's Github, so maybe the Discord engineers were unaware of said work or were unhappy with the pace of progress.
Kind of miffed at myself for missing that comment, since it directly addresses my question, but at least it's answered now.
[0]: https://old.reddit.com/r/programming/comments/eyuebc/why_dis...
It seems during that gap in time, the Go runtime team happened to solve their problem, which was GA in Go 1.12 in Feb 2019, which was prior to them doing the re-write in Rust, at least as far as I was able to follow.
One imprecise quote on timing of the re-write: [1]
This blog post perhaps is a bit "after the fact" we had made the switch over mid 2019
It's not crazy for someone to put aside a problem for a while, and then upon returning to the problem some time later decide to go a different route without re-exploring prior solutions, if that is what happened.
All that said, I might have misunderstood the timing, and I'm trying to avoid going back over all the various forums they commented in to find a better quote on timing ;-)
The situation you describe is certainly plausible, though. Shame that the chances of finding out what actually happened are quite low at this point.