Etcd 3.4
kubernetes.io
kubernetes.io
Further out, a case can be made that the whole business of reconciliation is one of event stream processing. My gut feeling for some time has been that the intersection of event stream and bitemporalism[1][2] makes a lot of the problems that motivate consensus protocols approximately moot. Done properly it might be possible to do away with the need for a master for a lot of things, which would improve scalability, failure tolerance and security.
[1] or Johnston's tritemporalism, I don't yet feel confident enough in my understanding to say.
[2] I am not alone and probably not the first to think so. For an example of a streaming-oriented bitemporal store, see Crux: https://juxt.pro/crux/index.html
It seems like a follower could pull the initial snapshot off of another follower to start instead?
The whole architecture is designed such that the etcd storage backend could be swapped out completely and the only thing that would care is the api servers. Much like the transition from etcdv2 over http with json objects to grpc with etcdv3 and protobufs.
You can also create alternative implementations of kubelet, kube-proxy, the scheduler, the controller-manager, etc because they all access the data via the api server's well-defined public facing API and anyone using the API can easily watch objects in any programming language using the same semantics as the rest of the API. It also works from browsers, etc.
Additionally, kubernetes supports RBAC for nodes themselves such that they can only see updates for objects related to pods running on them - you wouldn't want any node to be able to watch all secrets in the cluster needlessly.
Overall, I think we'd lose a lot if kubernetes switched to having all of its components access the data store directly. Every operator ultimately needs the same things the controller-manager, scheduler and the other kubernetes components need
I didn't ask about accessing etcd directly, i just wonder if http is the best transport for this case.
I'm glad this has been fixed, considering that one of the use cases for using partition-tolerant data stores is to tolerate partitions.
Cloud Foundry used earlier versions of etcd and this category of problem was the leading cause of severe outages. To the point that several years of effort were invested to tear it out of everything and replace it with bog-ordinary RDBMSes.
Disclosure: I work for Pivotal, we did a lot of that work, but I wasn't at the front line. Just watching from a safe distance.