In contrast, when a kubernetes controller or operator crashes, it can be expected to continue seamlessly where it left off.
It is easier to write kubernetes controllers that are able to continue seamlessly then to write terraform providers that do so, because of the granularity of the persistence of the state machine. Terraform locks the remote state, then applies all resources in the current root module, then unlocks the remote state again. In contrast, kubernetes operators can granularly update individual objects after each API call that is performed.
EDIT: Alright guys, good points all around. Mainly what I meant is the ratio of useful of Terraform vs its shortcomings. Not hating on the Cluster API pattern, but imho better to stick to standardized approach.
Besides that, the control-loop pattern, which the parent comment describes is a very sound design, which is employed not only in kubernetes and has nothing to do with "Optimizing based on a lottery event".
Is it though? GP claims to have been using terraform for years.
I have been using terraform for some time too, and from what i've seen well written providers will error out in a safe and controlled manner instead of making the whole thing crash.
> Even in "mature" ones as aws/gcp.
Can't comment for the gcp provider, but I've worked at companies where the terraform provider is used quite extensively and really can't recall a real crash that created serious problems. And i've been also using the openstack provider, again, quite extensively, with not much troubles.
So in conclusion I think that it's your argument that's really short-sighted, because it's based on FUD and situations that are theoretically possible but sufficiently remote in the real world that they can be ignored.
I'm not sure which part of my comment was "FUD" - I never said that you shouldn't use terraform, I was pointing out that issues exist, always. I think of it as the nature of the work we're doing (I say we, as I imagine you are also in the SWE field based on your comments). Show me any piece of software without bugs, especially as complex as a tf provider, and I'm buying you whatever you want.
[1] https://github.com/hashicorp/terraform-provider-aws/issues?q...
It's extremely painful and delicate to do and it may not happen frequently, but in my experience it happens frequently enough when I'm dealing with a lot of terraform.
That said, I agree in general that this wouldn't be the main reason to use the Cluster API over Terraform.
I think it depends on how often you "win" the lottery. If a lottery event happens once or twice a year and it's fairly hard to resolve, we classify them as land mines (or sea mines for the more maritime oriented people among us). These issues get a lower priority, but we'll try to work through them to avoid emergencies or lower their impact. So, maybe not optimizing, but increase resilience against lottery events.
Source: was a core developer of Terraform and some of the largest providers for many years.
Terraform isn't a wrong decision, just that CNCF now has a more stable API/CLIs you can work with, with which Terraform also needs to integrate.
Unsurprising because all the Hashicorp stuff is from a single vendor, and there is the risk of lock-in.
* simplified tooling - you can use all the same tools you already use to manage your k8s configs
* more standardized setup for multi-cloud/onprem
* it’s based on vanilla kubeadm and cloud-init/systemd images - debugging problems is easier bc it’s all pretty standard, battle-tested tech
When it works, it's nice to be able to deal with one templating language and one system to manage all of the elements of your deployments. Deploying a single helm chart is a nicer developer experience than managing both a helm chart and a terraform plan.
if anything that would be because the kubernetes operators implementing the Cluster API for your specific cloud provider would be implemented in a continuously running control loop, thus picking up changes... whereas with terraform you would have to run terraform from time to time.
edit: plus of course all the other features coming from kuernetes like mutating admission, validation, rbac (eg: only allow some people to create clusters or allow people to create clusters without granting access to the underlying cloud provider) etc etc...