https://blog.roblox.com/2022/01/roblox-return-to-service-10-...
TL;DR The outage was caused by (a) they enabled a new streaming feature in Consul under unusualy high read-and-write load, (b) the load conditions triggered a pathological issue in the third-party BoltDB system upon which Consul relies and (c) all of that was exacerbated by having one consul cluster supporting multiple workloads.
I'd use quotes to catalog this as a "horror" story, because this was clearly a very specific issue triggered by a specific and complex set of circumstances.
The blogpost also mentions that Roblox worked closely with Hashicorp engineers to mitigate the issue, and work towards structural solutions; the post also affirms their choice to manage their infra themselves rather then moving into a public cloud solution.
Sure, Kubernetes covers loads of territory. But there definitely are niches where products like Consul & Nomad do add value.
Enjoy spending the rest of your life trying to get etcd to cooperate. If you think operating Kubernetes at scale is a cakewalk, you don't have the scale problems you think you do.
I'll take consul over etcd ten million times out of ten.
If deploying on kubernetes, you now have all the problems that come with kubernetes, plus additional problems of trying to get consul working.
I spent a week just trying to stand up a federated multi-datacenter deployment of consul on EKS before my company decided it was too much hassle
I was saying that anyone who operates large scale Kubernetes knows that you will forever be dealing with tuning etcd and fighting to keep etcd alive. It's an underpinning service of Kubernetes.
Roblox's outage was related to the intricacies of running consul and mistakes that they made.
The point I was making was that I would rather, at this scale of operation, be running Nomad and optionally Consul and optionally dealing with the intricacies of Consul than running Kubernetes and being _forced_ to deal with what a miserable pain in the ass etcd is.
I was saying that running Kubernetes is _not_ the obvious choice -- at least once you're at 10^5+ systems.
The hyperbole applied to the description of the harm done to Roblox here I'm just going to ignore. Roblox is still popular. The company still has the reputation of a rock-solid engineering department. Roblox's stock isn't doing much differently than the rest of technology companies on the market and the "damage" you mention is vastly overstated anyway. In fact, Roblox's stock literally had a huge rally after and PEAKED within two weeks after the outage.