Zookeeper does expect a single “host” for every server but it runs perfectly fine with docker. But that’s no different from consul or etcd.
Zookeeper does expect a single “host” for every server but it runs perfectly fine with docker. But that’s no different from consul or etcd.
Should it not be the case that one of the responsibilities of the quorum is to vote new members in or out? I mean, as a first-class feature of the system.
The main problem I see is that in most consensus systems, any member can be nominated as leader. But if the member was inducted during a partition event - which one could do if the partition were long duration - then nominations for new leaders will go out. What happens if the new member gets elected? How do the partitioned machines find that leader when they return?
So the process of joining the quorum would have to be incremental. Because demanding a unanimous vote to add a machine means you can never replace a dead one (except by impersonating it)
Indeed. What you're describing is one of the main motivations for Raft. Paxos showed that distributed consensus was mathematically sound, but did little to guide implementors in actually building such a system. Raft is not a fundamentally new consensus algorithm; just an incremental improvement that formalizes many of the improvements that you needed to make to Paxos anyway, like membership changes, log compaction, and multi-decree support from the get go.
If you're interested, the Raft paper is quite readable and goes into this in detail. [0]
> The main problem I see is that in most consensus systems, any member can be nominated as leader. But if the member was inducted during a partition event - which one could do if the partition were long duration - then nominations for new leaders will go out. What happens if the new member gets elected? How do the partitioned machines find that leader when they return? How do the partitioned machines find that leader when they return?
This isn't actually the tricky bit, as it turns out. Communication flows from the leader to the other nodes, so the leader, even if it's a newly-inducted node, will initiate the connections to the partitioned nodes when the partition resolves. (The leader necessarily knows the addresses of the partitioned nodes, because the leader knows about all committed entries, and the identities of all the nodes in the cluster are committed into the Raft log.)
Again, the Raft paper does a great job explaining cluster mebership changes—much better than I can!
Goes to show you should always go back to the source at some point, even if the third party descriptions have better facility.
Dynamic reconfig in 3.5 addresses the "restarting every zookeeper instance" problem. [0] You stand up an initial quorum with seed config, then tie in new servers with "reconfig -add". Not sure how well it would tie into cloudy autoscaling stuff though. I wouldn't start there myself.
A much bigger pain IMO is the handling of DNS in the official Java ZK client earlier than 3.4.13/3.5.5 (and by association, Curator, ZkClient, etc.). [1] The former was released mid 2018 and the latter this year, so tons of stuff out there that just won't find a host if IPs change. If you "own" all the clients it's maybe not a problem, but if you've got a lot of services owned by a ton of teams it's ... challenging.
Even with the fix for ZOOKEEPER-2184 in place I'm pretty sure DNS lookups are only retried if a connect fails, so there's still the issue of IPs "swapping" unexpectedly at the wrong time in cloud environments which can lead to a ZK server in cluster A talking to a ZK server in cluster B (or worse: clients of cluster A talking to cluster B mistakenly thinking that they're talking to cluster A). I'm sure this problem's not unique to ZK though.
Authentication helps prevent the worst-case scenarios, but I'm not sure if it helps from an uptime perspective.
TL;DR: ZK in the cloud can get messy (even if you play it relatively "safe").
[0] https://zookeeper.apache.org/doc/r3.5.5/zookeeperReconfig.ht... [1] https://issues.apache.org/jira/browse/ZOOKEEPER-2184