Did the custom resource controllers pan out well?
Did the custom resource controllers pan out well?
Custom resources had/have some growing pains but I think worked out pretty well. It can be very very hard to test and distribute them though. As someone who maintained a controller with paid support, your test matrix gets pretty large pretty fast accounting for different k8s versions, different hosted versions (GKE, AKS, etc), different distros (openshift, rancher, etc). And that's before you even get into specific configurations like pod security policies, can the control plane communicate with the data plane, is there a service mesh.
Resource versioning is hard to get right. Once a resource type is v1, it becomes difficult to extend it. You can't add a beta field to it easily. Revising schema can be hard since anything more than a no op conversion between versions requires a webhook, which requires a certificate chain, and while cert-manager is popular it is not ubiquitous and regularly has breaking changes. Webhook setup issues made up a large portion of our support requests.
As far as the general "reconciliation loop" architecture goes, you end up with something similar in most orchestration systems I've worked on, or you wish that you did. So overall I think that worked out well. Getting it right can be hard, but I think that's the nature of the beast.
It was never reactive, at least in the sense of the Reactive Manifesto. There is no system of backpressure.
> Did the custom resource controllers pan out well?
I think it will prove to be useful for a handful of uses, but that in general, it will wind up like Wordpress and Jenkins plugins. Powerful, popular, but a guaranteed mess.
Kubernetes was not originally designed to accommodate CRDs. There's no concept of tenancy, only things you can cobble together into tenancy-ish-alike-kinda shapes. There's RBAC, but it was designed before CRDs and means that poorly designed custom controllers can be attack vectors.
I think CRDs are useful and necessary[0], but that they are greatly overused and have expensive design flaws. Unfortunately, I don't have the wits to join into the multi-billion dollar market of "papering over things Kubernetes doesn't do very well but which can't now be changed".
[0] I mean, I wrote a book about Knative, which is a purely CRD-based architecture.
Could you elaborate on the "design flaws" part? Do you feel that the CRD model itself has design flaws or are you referring to specific project's CRDs?
——-
The usage patterns of native k8s types and the implications those patterns have on the scalability and reliability of etcd and the apiserver are relatively well-understood. CRDs can be a wild-card, though, and afaik testing efforts thus far have not investigated worst-case usage of CRD-based applications.
As commonly deployed, CRDs are served from the same apiservers and etcd cluster that serves the native types for a k8s cluster. That can result in contention between the CRDs supporting 3rd party additions to a cluster and the native types critical to the health of a cluster. This kind of contention has the potential to bring a cluster to its knees.
Efforts like priority and fairness seek to ensure that the apiserver can prioritize at the level of the API call. But that won’t prevent watch caches from OOM’ing the apiserver if excessive numbers of CRDs are present. The judicious use of quotas could head off the creation of an excessive number of objects, but it’s not just count that matters - the size of each resource is also a factor.
In theory, CRDs could be isolated from native types by serving them from an aggregated apiserver backed by a separate etcd cluster. afaik this not a supported configuration today, and even if it were the additional resources required to support it (especially the separate etcd cluster) may be prohibitive for many use cases.
You can actually nominate particular types be stored in particular etcd servers -- GKE does this to put Events into a separate etcd from everything else.
However, it still has problems. Firstly, you can only define it for inbuilt types. Secondly, it's common for different objects to cross reference each other through objectRefs and the like, which behave badly when you effectively perform a join in the API server over multiple etcds.
Interesting. Is this documented anywhere?
What I said was slightly wrong. It's not that you nominate Kinds, it's that you nominate which etcd servers get which etcd key paths. You can essentially work it out because the path structure is consistent.
>"But operating with CRDs at scale is a different story and suggests careful testing with the specific applications involved."
Do you mean the number of different CRDs deployed here or just the number of custom resources created? Or is it the same concern with either? I'd be curious what you are defining as "scale" as well?
Scalability is relative, and depends on many factors including but not limited to:
- the resources available on the hosts running apiservers and etcd members
- the number and size of resources (custom and native) that controllers will maintain
Relatively speaking, a cluster of a given size might be perfectly capable of handling on the order of many thousands of resources . Push that an order of magnitude and the overhead of serving LIST calls - marshaling json from etcd to golang structs for apimachinery and back again for sending over the wire - could exhaust an apiserver’s memory allocation. And since the impact of resources is cumulative, any one application relying on lots of CRDs might not destabilize a cluster on its own but might well contribute to an unhealthy cluster when running alongside similarly CRD-heavy applications.
The key takeaway is that the kube api is best thought of as a specialized operational store rather than a general-purpose database. Anyone wanting to rely on CRDs at non-trivial scale would be well-advised to test carefully.
Controllers can be scoped to a namespace without requiring cluster-level permissions, no?
It means every controller for a CRD winds up installing another webhook. And having to test a variety of orderings. It's hard to get right.
> Controllers can be scoped to a namespace without requiring cluster-level permissions, no?
The difficulty is that Kubernetes RBAC is good at expressing rules "Role 'foo' can perform operation 'list' on kind 'Deployment'". But it's less capable of saying something like "Role 'foo' can perform operation 'list' on kind 'Deployment' which were created from kind 'CoolerDeployment'". It's also hard to delegate something, along the lines of "Role 'foo' can delegate ('create' over kind 'Deployment') within namespace 'bar'".
I think dissatisfaction with PodSecurityPolicy will cause Rego to worm its way into the core architecture over the coming years. It'll then eventually crowd out RBAC because you can impose (sort of, more or less) arbitrary rules. But none will dare call it ABAC.