Cloud infrastructure should be immutable
stroobants.dev
stroobants.dev
Oh yes we will indeed just redeploy that container. Oh well no we can’t because the PV is stuck.
Oh yes we’ll just redeploy all the SQS queues. Oh no wait they have data in them.
Immutability, IMHO, is a terrible idea for infrastructure as a general purpose rule. Making it more complicated for developers to deploy systems, and requiring deployments to go through many steps (as immutability typically requires a full blue / green deploy), means MTTR will be longer.
Also, trying to fix a problem in situ takes longer (because you don’t have the tools for it).
Lastly, immutability is a lie. If you’re using something like v8, over time, the JIT and GC may change behavior. The cloud vendor may deploy new software.
If your software can be deployed as a fully self contained, stateless application, perhaps you can ignore these problems, especially by constantly recycling infrastructure. Now, the real problems begin to occur when you stop recycling - that’s when really interesting longevity problems appear.
/rant
I personally see immutability as a sign of organizational immaturity around system administration, if you're concerned about drift on production systems you don't have the right config management and RBAC in place. It's never more efficient to replace a fleet vs patching it in place, especially from a time standpoint.
Deploying new code: publish a new container, scale up green, switch, whack blue at your leisure.
Patching: publish a new AMI, scale up green, switch, whack blue at your leisure.
1. Every night we build a new freshly patched Ubuntu image.
2. Three times a week we run a deploy that rebuilds our entire infrastructure in place with the latest Ubuntu image as the base [1].
3. Once the deploy runs we never run apt install or apt update again.
So rather than
server.update()
we do server = server.update()
[1] This sounds fancy but I promise it's stupid simple, all the complexity is in the auto scaling groups and the load balancers which we don't manage. And it's the exact same procedure to do code deploys which we do multiple times a day.> When an error occurs, operations can just redeploy
I mean no. If there is an error you should figure out what the problem is first.
Also if operations just blindly redeploy when things go wrong you have some large culture issues that need to be solved first.
> Operations can quickly return to a previous state
I kinda agree. However the hardest thing to fix normally is the state that was fucked up by what ever outage that happened. Assuming that you have the correct mechanism for inter-service comms, its really simple to kill/restart services, so long as you have no data in flight.
Whats not easy to fix is if the state that was stored is causing the crash. Those are normally the big problems. Bad messages/config causing rolling outages.
> Your operations team can adopt a workflow > [..]You were dropped with a bunch of co-workers and all of them did their own thing[..]
I mean thats is a massive problem right there. There is either the right way to deploy things, or there isn't. That's a culture problem and needs to be fixed, and fixing it is very very hard.
On the hole I broadly agree that Prod should be >95% immutable, and nowadays its fairly simple to achieve.
One thing I would recommend is running "prod" in two regions with some sort of health checked loadbalancer infront. This allows testing of IaC stuff in prod without bringing the entire service to a grinding halt when stuff fails. Yes managing state can be much harder, and its important to partition your state by region so you are not hampered by cross region syncing.
But the advantages are that you are much better placed to survive an outage, or resource crunch. But most importantly it allows you to completely kill and rebuild a "prod" instance from almost scratch. This is key for DR preparedness.
Figuring out what the problem is may take time. If a problem is impacting a customer’s experience, I’d rather recover faster and then pull metadata later to try and understand the issue.
How do you know it'll fix the problem, if you don't know the cause?
In my experience, about 50% of the time to fix an incident is cleaning up the mess you made performing a fix based on an hypothesis that turns out to be wrong. That means longer downtime.
A classic example of this is when a 60 disk raid unit crashed when performing a rebuild under load. This happened to my coworker when they were doing the daily replacing a disk task. They put in the disks, shut the draw, pressed the button that changes the activity lights to see that everything was working, and the thing froze.
They rightly shat their pants. They panicked, rebooted the array and attached fileserver. A load of data was lost (the file servers had at the time the maximum amount of ram you could shove in a standard 2u two proc intel server).
When it happened again, another coworker waited for the raid enclosure to reboot. Low and behold the SAS links came back and we were able to do an online fsck inside 3 minutes. no data lost.
If your facing an outage, breathing space to gather thoughts is time well invested. Blind hacking, or indeed decision paralysis is damaging
First thing you want to do (after confirming the issue and communicating) is to "fly the plane" and try to contain/mitigate the issue before debugging for a RC.
Unless there was data corruption or changes caused by the first issue.
>When a disaster occurs, operations can just deploy in another region
Except all the data that is in the old region in databases, ecr repos, s3, etc, etc.
>Operations can quickly return to a previous state
Unless something like a database migration was done.
Yes, we want to have declarative configurations and idempotent "terraform apply".
Yes we want to have containers and subsystems with no persistent state, and careful architecture for persistent storage volumes and database services.
And please remember to say "prevent_destroy", "ignore_changes" etc for that persistent data where appropriate. And remember that the automation that creates your system in one command will be able to destroy it in one command, if you say so.
lots of resources are region-bounded. it's just not that easy.
Also there’s contention whenever that happens. Last AWS outage we experienced, entire availability zones had no capacity available. If they did it took forever to spin anything up and migrate volume snapshots over as well as everyone else was doing the same thing at the same time.
To find a root cause, I need to look at the broken thing. If you hit restart, the only thing I can guarantee is that the fault will occur again.
This is newbie hipster webdeveloper BS.
Immutable infrastructure is great for stateless services, to ensure you know exactly what versions are used and to avoid people from making changes on the server. For databases and other stateful services, such as message brokers, it's less ideal.
All the tech in the world can't mend a broken process or org chart!
BTW - drift detection is a harder problem than it seems and has a lot challenges running this accurately at scale. I wrote about this here - https://www.cloudquery.io/blog/announcing-cloudquery-terrafo...
Full disclosure - Im the founder of CloudQuery
Azure reached out though, and in so many words asked: wtaf are y’all doing and can you please stop?