This is pretty much the framework I teach admins at work:
- What impact does the change have?
- How do you get out of that change, and how long does that take?
- And what is your confidence into the change, and the bail-out plan?
And honestly, if you have a high-confidence, fast bailout plan, you can be downright brazen/#yolo about changes. We've recently had to update a central and critical IDP, but we eventually realized: We have the old docker images, and it has a 200MB sized DB. We can dump + restore that in 2 minutes. So if the upgrade goes wrong, we have high confidence to rollback in like 5 minutes. At such a point.. why not just go with it?
Similar things are developing with Postgres upgrades. Setup 1 leader + 3 replicas, upgrade 2 replicas, failover, see how much explodes and at worst, fall back. If we can test beforehand, alright.
Other teams plan complicated upgrades requiring coordinated actions of 6 other teams. And like 3 know how to possibly take back that change? And like 4 know what to actually do? Ugh, this ended up in a fun weekend.