I feel for the admin here. This seems equivalent to an accidental `sudo rm -rf /`.
I feel for the admin here. This seems equivalent to an accidental `sudo rm -rf /`.
...does absence signify deletion?
Some projects decide yes, if the YAML specifying your resource is absent on the next reconciliation, then delete the corresponding K8s resources.
Then some projects go with "no, deletion must be explicit, set this flag in your YAML".
But the last approach doesn't work as nicely with GitOps, as you have to do a two stage workflow to delete - first set the delete flag, then once the resources are deleted, then delete the YAML.
But I'm okay with deletion being harder to do if it makes it harder to well, accidentally delete all your stuff.
Because I've never met a K8s workflow where resources were commonly deleted, usually it's creation and update 90% of the time.
But I can understand it'd be annoying if you ended up with dangling resources if you missed an explicit delete flag.
I also argue that guardrails are equally important for imperative migration tools as well, but more often they're lacking or half-baked, which gives a false sense of safety. For example, down/reverse migrations are a very common landmine for human error. Order-of-operations problems also happen frequently when imperative tools are used by large dev teams, resulting in subtle schema drift between environments when there's disagreement between lexicographic migration file order, git history order, and the actual migration application order on each DB.
I have all the sympathy for the author. I’m sure they don’t feel good right now, but I hope they continue contributing. They’ve learned a hard lesson, but putting it back into their practice is the only way to make it count.
You can use an admission controller to add guard rails, but I'd hope my GitOps operator also offered those guard rails.
[1] https://argo-cd.readthedocs.io/en/stable/user-guide/auto_syn...
[2] https://kubernetes.io/docs/concepts/storage/persistent-volum...
I'm hard-pressed to think of a scenario where a single sysadmin in an 'old-fashioned' enterprise datacenter could do as much damage with, maybe not exactly a one-liner but a lightweight change instruction.
The risk-reward calculation seems completely bananas to me.
So deletions are explicit and a mistake is just un-abandoning.
We're planning to streamline our run approvals a bit, because approving every DNS record addition results in some approval fatigue, but resource deletion will most likely always require human approval.