The v3 API of ectd seems more powerful because it supports multi-key transactions but isn't there yet.
Other than that I always had a slight feeling that Consul might be a bit more stable/reliable but I have no hard facts.
The v3 API of ectd seems more powerful because it supports multi-key transactions but isn't there yet.
Other than that I always had a slight feeling that Consul might be a bit more stable/reliable but I have no hard facts.
- We have a functional testing cluster that is constantly simulating various types of failures. You can find some information on that here: https://coreos.com/blog/new-functional-testing-in-etcd/
- The team is works to ensure the raft library that we use (and that we now share with cockroachdb) has unit tests that handle all of the edge cases described in the raft paper. See this paper as an example: https://github.com/coreos/etcd/blob/master/raft/raft_paper_t...
- We have been running etcd in production for 1.5 years for a bootstrapping service called discovery.etcd.io and have focused on how to handle higher and higher read/write loads there. For example see the improvements we made in etcd 2.1 with async snapshots: https://twitter.com/BrandonPhilips/status/639843455141740544
- We have also focused on clear guides on how to operate etcd. All of this testing and performance work is of no use if people can't operate the system correctly: https://coreos.com/etcd/docs/latest/
Of course if you have particular issues that you have encountered we would love to help: https://github.com/coreos/etcd#contact
Are you using his jepesen testing tool as part of your testing for failure scenarios? If not, why?
If they rewrote that library and CockroachDB (whos engineers I respect) uses it, then I guess my fears are unwarranted.
We don't have jepesen setup as a testing tool because it is really hard to make it a reliable false positive free system. Plus, the languages it is written in make it hard for us to hack on it.
Instead what we have done is worked hard to build a functional testing suite to find non-algorithm issues (weird exhaustion issues, behavior under real disks, etc). And then have deep, fast, deterministic tests of the core raft algorithms.
I'd disagree with the stable/reliable bit; or, at least, I'd note that CoreOS ships with etcd, and fleet (the CoreOS job manager, at least until Kubernetes obviates the need for it...) depends on etcd to function correctly.
The upshot of that is, if you're running a stable release of CoreOS, you have a stable etcd, by definition - or you have a broken cluster that you can't submit jobs to, which would be bad.
From the outside looking in, to me Consul seems to provide more features; I don't doubt that it's awesome, but etcd ships with CoreOS, and when I spin up a new cluster via Cloudformation it's up and waiting for me as soon as I log in. This is a huge selling point, as I am lazy.
I've used Zookeeper in the past for a dosgi monstrosity and it was a never-ending source of pain. I haven't looked at Consul, but since I already have a good grasp on how Etcd works, I probably won't bother with it.
- locksmith: a scheduler for host reboots in a cluster. Designed to ensure a cluster can do OS upgrades unattended.
- skydns: a DNS server built on top of etcd.
- confd: a configuration file templating system designed to watch changes and rewrite configuration on disk.
- vulcand: a HTTP load balancer with rate limiting, and dynamic balancing algos.
- kubernetes: a system to manage clusters of containers which backs its service discovery, scheduling, and election with etcd.
One interesting thing is that kubernetes handles load balancing, DNS and configuration packaged together which many people enjoy using. While other people like to just have DNS or configuration so there are tools that focus on just that to tie together existing systems.
GET /v2/keys/mykey?quorum=trueThe link to the network partition I'm talking about is buried in this presentation http://thesecretlivesofdata.com/raft/
This assumes all reads and writes go through Raft.