If it works for me it works for everyone.
You never patch or upgrade immutable infrastructure. You just replace what you've got with a new VM or container. Containers being preferred because they can be started & stopped near instantaneously and there's nothing like a virtual BIOS that could have different configurations like with VMs.
You don't stand up a VM or container then "log in to configure it". Once the VM is "up" that's it. You're done. At that point you just need to point your load balancers/DNS at the new stuff then take down the old stuff.
One interesting aspect of immutable infrastructure such as this is that it is completely incompatible with loads of existing security policies and what would have been considered "best practices" just a few years ago. For example, you might have a security policy that states that everything must be scanned within 30 days for malware/out-of-date packages/whatever. Yet with immutable infrastructure your hosts or containers may only be up for a few days before being replaced!
So when your security team freaks out because none of your hosts/containers are showing up in their systems you'll have a lot of explaining to do =D
"We need to scan your hosts so we can ensure that you're installing security patches."
"We don't do that."
"You don't install security patches?!?"
"Yeah, well, you see..."
Trust me when I say that trying to explain how it all works and why it's more secure than old school deployments is not easy!
FROM ubuntu
WORKDIR ${foo} # WORKDIR /bar
RUN developer_script.sh
So now with each new container update you'd need someone from the security team to audit/review `developer_script.sh`. It's kind of pointless if your goal is fast deployments.If you just need to make sure your developers don't make a mistake in terms of securely configuring their containers (and making sure to always use the latest software) then you simply scan them before the canary stage. The problem there is, "what are you looking for?"
Also, there exists only one tool to scan Docker containers (OpenSCAP) and if your security team doesn't like it, well, you're screwed: https://github.com/OpenSCAP/container-compliance
The other problem is that OpenSCAP only checks the container's packages for compliance. It doesn't actually scan the container's filesystem for things like JREs and bundled libs. So if you're using Docker best practices by keeping your images as minimal as possible you may not even have a package tool inside your containers. In that case how do you check for things like out-of-date versions of Java?
Another problem is you assume there's some modicum of control over what's inside the containers before they're deployed. We're in charge of creating the infrastructure for running containers with the promise that end users (who would be various application teams) can make their own containers (or at least their own Dockerfiles) and deploy them on our infrastructure.
We can work with them to help develop, say, a Kubernetes pod config for their app but as far as what their app is or what gets bundled with it we'll have no knowledge after the first successful deployment.
tl;dr I assume there's some control because there should be.
> "We need to scan your hosts so we can ensure that you're installing security patches."
> "We don't do that."
> "You don't install security patches?!?"
> "Yeah, well, you see..."
If you use a tool like zypper-docker, you can create a new image quickly that applies just the security patches. Currently it only works for SUSE-derived containers but we're planning on making it distribution agnostic.
There's also some work we're doing on connecting Docker containers running on SUSE enterprises to connect with SUSE Manager (aka spacewalk), so they would show up in their systems.
You can find most of this stuff in github.com/SUSE.
Our Docker registry pulls down the latest images (that we use with our Dockerfiles) multiple times daily. So if the author of the Dockerfile just puts "FROM whatever:latest" they are guaranteed to have all the latest/patched software whenever they "docker build".
The point of zypper-docker is to quickly roll out security fixes with essentially no downtime. zypper-docker also allows you to quickly check the health of all of the images and containers on your servers, so you can figure out which ones need updating. Also, zypper-docker allows you to apply RPM patches (something that's quite crucial for enterprises).
So no, it's not pointless. In fact it solves a problem that not many people have been working on solving: updating containers.
How are you going to know if the latest version of the upstream image breaks your package if you don't try it? E whole point of continuous integration is that there are no surprises when it comes time to push to production.
If ":latest" breaks your image you had better know:
A) Immediately.
B) What went wrong ASAP.
The integrated testing (that you're supposed to use with Docker best practices) should reveal if there's a problem and if it doesn't you'll still catch it during QA testing of your image.
The fact that ":latest" breaks your image should be nothing more than a few minutes to days of troubleshooting. While you're doing that your existing production images will keep on chugging away.
The only rule that you must not break is that you have the next release of your image out the door, ready for production within 30 days. Why? Because that's the maximum (hard-coded) lifetime we (and you should too) allow images to stay running.
Would you fix a software bug by editing the code on a running server and tell yourself that you will add it to the repository later? Of course not. You would end up with a running instance of the code that is impossible to replicate.
Immutable infrastructure applies that same idea to running services.
Absolutely yes in the right circumstances (have you seen how this website works?).
Not every service needs to be written as "scale out to a billion nodes" architecture with eight layers of checks and balances between idea to production.
We can't globally say "everybody must use a 16 step verifiable app development pipeline" when some people run 3 servers and others run 3 million.
(Plus, not all services are stateless! Redeploying stateful servers is painful—you'll be killing user experience. Updating code live on a running server while maintaining internal continuity of state can maintain sessions/caches/game-state without annoying users by kicking out of their current flows.)
I maintain a personal server that hosts a variety of VMs for things like mail, static site hosting, nameservice, etc. Last time I rebuilt it, I thought, hey, I'll try doing it with these interesting tools. The immutable-infrastructure approach turns out to be pretty heavy in that case. A lot of times it feels like I have to perform a triple bank shot just to make a minor configuration change.
That said, I think the one-off server, carefully maintained by one person, is becoming very rare. When I build things that other people will work on, I think immutable infrastructure and the cattle-not-pets approach is the only responsible way to go. Even if traffic volume won't be huge, I think the clarity and ease of debugging you get is vital when it doesn't all live in one wizard's head.
I agree that stateful servers are a challenge in this context. But they're a challenge regardless. Having that one thing we're afraid to upgrade or restart can cause all sorts of development and business process issues. That might have been a worthwhile tradeoff 10 years ago when you had to buy all your hardware and you needed ops people to do a lot of stuff manually. Servers go down, and we might as well accept that from the beginning.
For example: update the code in your source control repository, and then build the new version of the software.
It's not, and that's coming from someone who does devops for thousands of virtual machines in production. I don't have an hour to burn sometimes to update a repo, build new AMIs, roll them into production, roll the old ones out, watch logs to ensure canaries pass acceptance tests and that production traffic to new instances aren't erring out, all for minor changes (example: nginx mime type change).
Automatic (and safe, obviously) production cutover is still kind of a holy grail, but it's definitely doable, and being done as we speak by a number of pretty neat companies.
The image the EC2 instance is started from is immutable while it is running. Of course the image gets updated regularly by devops with security fixes, but the running production instance(s) is/are never changed on the fly. Instead it is completely redeployed.
It's the diff between a new build (of the "same code") for each environment versus building up the VM one and deploy that.