As a person with no background in distributed systems, I am wondering why people choose Consul over alternatives. Are there features that etcd doesn't offer?
As a person with no background in distributed systems, I am wondering why people choose Consul over alternatives. Are there features that etcd doesn't offer?
I don't believe etcd would have been any better for us, though. Centralized service discovery that runs through raft consensus doesn't make a lot of sense for the things we need to do. And when I've had etcd blow up on me in the past, it's been similarly painful to recover from.
Most people don't even know that the Kubernetes control plane by default has a hard limit on etcd size. It used to be 2GB, not sure what it is now.
I have seen very few strongly consistent distributed KV store that scales beyond 10GB+
However, related to that, for big-time clusters (q.v. https://news.ycombinator.com/item?id=35174655 and https://news.ycombinator.com/item?id=25907312) one should without question move events over into their own etcd cluster: https://openai.com/research/scaling-kubernetes-to-2500-nodes...
There’s also max object size of 1MB on the apiserserver side I believe
The single-group raft is the hard limit.
Heh, that kubebrain TODO is some "oh, really?"
* Guarantee consistence in critical cases
but I give them huge props for calling out Jepsen
AFAIK doesn't Consul also use Raft?
I think I understand how you're using it and curious if you've considered how AWS STS API manages their cross region syncing gets solved.
If you want apps to discover each other and be able to communicate effortlessly, even across datacenters, Consul, in theory, enables this.
I say in theory because I couldn't get federated Consul actually working.
I used consul for a clustered service once, it was worth it for bringup. but I when I had problems I just wrote one in a couple days since I'd done so several times before. and it didn't fail for all the years that product was running.