I think that the future of deployments means that it is closely integrated with your source code management. Every new push builds a container that can go through the following steps:
1. Deployment (server with only test traffic)
2. Post-deploy test (smoke test)
3. Canary (part of the traffic)
4. Live / multiregion deploy
5. Manual overrides http://techblog.netflix.com/2015/11/global-continuous-delive... "Spinnaker also provides cluster management capabilities and provides deep visibility into an application’s cloud footprint. Via Spinnaker’s application view, you can resize, delete, disable, and even manually deploy new server groups using strategies like Blue-Green (or Red-Black as we call it at Netflix). You can create, edit, and destroy load balancers as well as security groups."
Also see https://gitlab.com/gitlab-org/gitlab-ce/issues/3286
Are there any downsides to immutable infrastructure?
However, practices like those tend to be a crutch anyway.
* It can get complicated very quickly since it involves a lot of stages, systems, and loads of brand new tools where you're unlikely to be able to hire existing talent (you need people who can just read the docs on a complicated tool/architecture like Kubernetes and then start using it).
* Speaking of talent, you need developers who are reasonably well-versed in Linux systems administration. Could your developers write a shell script that can configure every little aspect of a Linux host from packaging to authentication? Your developers also need to understand some hard core topics of TCP/IP (e.g. anycast), firewalls (e.g. port mapping), and DNS (TTLs, views, SRV records, and more). They also need to be well-versed in security (e.g. the impact of running as root inside a container) and especially authentication (e.g. HMAC, securely storing shared secrets) and encryption (SSL, managing certificates, knowing which hashing algorithm to use)!
* Dealing with existing bureaucracy. Especially rules and policies in regards to security. For example, some people will think that containers and immutable VMs should be treated like regular hosts. That can mean waiting hours for an inventory system to acknowledge their existence before you can put them into production. Or having to wait 30 minutes after every container/VM is brought up so that a security tool can perform a scan even though each container/VM would be identical in every way and have already been scanned before production deployment. Sigh.
Of course, a lot of that depends on where you work. At my work right now the biggest challenge is the bureaucracy for certain. If I can get containers classified differently than VMs then all my policy/BAU problems will (likely) go away. If I can't get that then, really, there's no point to our efforts since there would be no advantage to end users (the consumers of what we're building).
I may be able to live with a pointless, half-hour-long security scan every time we bring up a new container but I absolutely cannot live with an 8-hour-long wait for a container to show up in our system-of-record (inventory system).
Our main issue is also security - as docker containers aren't very good at actually containing, from a security PoV, we're using SELinux MLS, assigning each container a unique SELinux level/category (fedora-atomic sets this up nicely for you if anyone wants to try).
This is great for security, but breaks so many things - lots of the cutting edge work with docker and related tech isn't done by people that care deeply about security, or really understand this sort of usecase, so we run into a huge amount of edgecases and bugs. We spend a lot of time fighting this.
It also makes having a really good build/deploy pipeline essential, as well as a good environment set up to reproduce and diagnose issues.
So if you have a container running the database for client A that means that when client B spins up their little malicious container they can directly connect to client A's DB.
That's one way that Docker containers can "leak" (sort of). Another way is that the filesystem for running containers is directly accessible from outside the container. That means that anyone with the power to start/stop containers has the power to read the filesystems for all other containers.
There used to be a problem with, "root inside the container is the same as root outside the container" but that was fixed with Docker 1.9 via the new UID mapping feature. So the filesystem and processes running as UID 0 inside the container will actually be some high UID in the parent OS. So unless you're running Red Hat's version of Docker (which is currently hard set to be Docker 1.8!) you can upgrade and solve that problem pretty quickly with a few tweaks.
I am only slightly surprised that there's still a lot of bugginess in this aspect, though I haven't had to deal with that level of security need on the infrastructure side in a number of years. I'd rather be doing app-dev over ops-dev.
Joyent's Triton looks cool: https://www.joyent.com/ as does CoreOS's Tectonic (built on top of k8s): https://tectonic.com/
Hadn't looked at Atlas... I changed jobs into a larger company, and don't really have to deal with the ops side ever. Though eagerly awaiting a blessed docker based solution to become available in the org... it'll mean that projects can be developed more readily without being limited to the Java based infrastructure in place now.
The current popular method is to simply query a service discovery resolver of some sort (e.g. https://docs.docker.com/swarm/discovery/) to figure out which server to query, say, the weather (very important in finance! haha =). That's great and necessary if your architecture spans across your internal cloud, Amazon, and cloud-provider-of-the-day but if you have a fixed set of data centers where you know your applications will always be running (say, because of legal reasons =) then you can just use anycast as long as you keep ports consistent.
For most microservices deployments I'd imagine they would want their stuff to be accessible via a global DNS name and if they're going to do that then they might as well take the next step and just standardize on a port as well. If you have a single (global) name and a single port for an API why do you need service discovery (for that)? You don't. You just need to make sure that when clients query your API that they don't wind up using that one server you've got on the other side of the world on a 56k ISDN line (sadly yes, that does happen).
Having said that we may end up using popular service discovery tools anyway because they're convenient and we can use them in horrible, do-as-I-say-not-as-I-do ways to hack legacy applications into containers =D
1. Deploy directly from GitLab to Kubernetes https://gitlab.com/gitlab-org/gitlab-ce/issues/14040 2. Add a container registry to GitLab itself https://gitlab.com/gitlab-org/gitlab-ce/issues/3299
Any comments or suggestions are welcome.
It's the diff between a new build (of the "same code") for each environment versus building up the VM one and deploy that.
The image the EC2 instance is started from is immutable while it is running. Of course the image gets updated regularly by devops with security fixes, but the running production instance(s) is/are never changed on the fly. Instead it is completely redeployed.
Would you fix a software bug by editing the code on a running server and tell yourself that you will add it to the repository later? Of course not. You would end up with a running instance of the code that is impossible to replicate.
Immutable infrastructure applies that same idea to running services.
Absolutely yes in the right circumstances (have you seen how this website works?).
Not every service needs to be written as "scale out to a billion nodes" architecture with eight layers of checks and balances between idea to production.
We can't globally say "everybody must use a 16 step verifiable app development pipeline" when some people run 3 servers and others run 3 million.
(Plus, not all services are stateless! Redeploying stateful servers is painful—you'll be killing user experience. Updating code live on a running server while maintaining internal continuity of state can maintain sessions/caches/game-state without annoying users by kicking out of their current flows.)
I maintain a personal server that hosts a variety of VMs for things like mail, static site hosting, nameservice, etc. Last time I rebuilt it, I thought, hey, I'll try doing it with these interesting tools. The immutable-infrastructure approach turns out to be pretty heavy in that case. A lot of times it feels like I have to perform a triple bank shot just to make a minor configuration change.
That said, I think the one-off server, carefully maintained by one person, is becoming very rare. When I build things that other people will work on, I think immutable infrastructure and the cattle-not-pets approach is the only responsible way to go. Even if traffic volume won't be huge, I think the clarity and ease of debugging you get is vital when it doesn't all live in one wizard's head.
I agree that stateful servers are a challenge in this context. But they're a challenge regardless. Having that one thing we're afraid to upgrade or restart can cause all sorts of development and business process issues. That might have been a worthwhile tradeoff 10 years ago when you had to buy all your hardware and you needed ops people to do a lot of stuff manually. Servers go down, and we might as well accept that from the beginning.
For example: update the code in your source control repository, and then build the new version of the software.
It's not, and that's coming from someone who does devops for thousands of virtual machines in production. I don't have an hour to burn sometimes to update a repo, build new AMIs, roll them into production, roll the old ones out, watch logs to ensure canaries pass acceptance tests and that production traffic to new instances aren't erring out, all for minor changes (example: nginx mime type change).
Automatic (and safe, obviously) production cutover is still kind of a holy grail, but it's definitely doable, and being done as we speak by a number of pretty neat companies.
If it works for me it works for everyone.
You never patch or upgrade immutable infrastructure. You just replace what you've got with a new VM or container. Containers being preferred because they can be started & stopped near instantaneously and there's nothing like a virtual BIOS that could have different configurations like with VMs.
You don't stand up a VM or container then "log in to configure it". Once the VM is "up" that's it. You're done. At that point you just need to point your load balancers/DNS at the new stuff then take down the old stuff.
One interesting aspect of immutable infrastructure such as this is that it is completely incompatible with loads of existing security policies and what would have been considered "best practices" just a few years ago. For example, you might have a security policy that states that everything must be scanned within 30 days for malware/out-of-date packages/whatever. Yet with immutable infrastructure your hosts or containers may only be up for a few days before being replaced!
So when your security team freaks out because none of your hosts/containers are showing up in their systems you'll have a lot of explaining to do =D
"We need to scan your hosts so we can ensure that you're installing security patches."
"We don't do that."
"You don't install security patches?!?"
"Yeah, well, you see..."
Trust me when I say that trying to explain how it all works and why it's more secure than old school deployments is not easy!
FROM ubuntu
WORKDIR ${foo} # WORKDIR /bar
RUN developer_script.sh
So now with each new container update you'd need someone from the security team to audit/review `developer_script.sh`. It's kind of pointless if your goal is fast deployments.If you just need to make sure your developers don't make a mistake in terms of securely configuring their containers (and making sure to always use the latest software) then you simply scan them before the canary stage. The problem there is, "what are you looking for?"
Also, there exists only one tool to scan Docker containers (OpenSCAP) and if your security team doesn't like it, well, you're screwed: https://github.com/OpenSCAP/container-compliance
The other problem is that OpenSCAP only checks the container's packages for compliance. It doesn't actually scan the container's filesystem for things like JREs and bundled libs. So if you're using Docker best practices by keeping your images as minimal as possible you may not even have a package tool inside your containers. In that case how do you check for things like out-of-date versions of Java?
Another problem is you assume there's some modicum of control over what's inside the containers before they're deployed. We're in charge of creating the infrastructure for running containers with the promise that end users (who would be various application teams) can make their own containers (or at least their own Dockerfiles) and deploy them on our infrastructure.
We can work with them to help develop, say, a Kubernetes pod config for their app but as far as what their app is or what gets bundled with it we'll have no knowledge after the first successful deployment.
tl;dr I assume there's some control because there should be.
> "We need to scan your hosts so we can ensure that you're installing security patches."
> "We don't do that."
> "You don't install security patches?!?"
> "Yeah, well, you see..."
If you use a tool like zypper-docker, you can create a new image quickly that applies just the security patches. Currently it only works for SUSE-derived containers but we're planning on making it distribution agnostic.
There's also some work we're doing on connecting Docker containers running on SUSE enterprises to connect with SUSE Manager (aka spacewalk), so they would show up in their systems.
You can find most of this stuff in github.com/SUSE.
Our Docker registry pulls down the latest images (that we use with our Dockerfiles) multiple times daily. So if the author of the Dockerfile just puts "FROM whatever:latest" they are guaranteed to have all the latest/patched software whenever they "docker build".
The point of zypper-docker is to quickly roll out security fixes with essentially no downtime. zypper-docker also allows you to quickly check the health of all of the images and containers on your servers, so you can figure out which ones need updating. Also, zypper-docker allows you to apply RPM patches (something that's quite crucial for enterprises).
So no, it's not pointless. In fact it solves a problem that not many people have been working on solving: updating containers.
How are you going to know if the latest version of the upstream image breaks your package if you don't try it? E whole point of continuous integration is that there are no surprises when it comes time to push to production.
If ":latest" breaks your image you had better know:
A) Immediately.
B) What went wrong ASAP.
The integrated testing (that you're supposed to use with Docker best practices) should reveal if there's a problem and if it doesn't you'll still catch it during QA testing of your image.
The fact that ":latest" breaks your image should be nothing more than a few minutes to days of troubleshooting. While you're doing that your existing production images will keep on chugging away.
The only rule that you must not break is that you have the next release of your image out the door, ready for production within 30 days. Why? Because that's the maximum (hard-coded) lifetime we (and you should too) allow images to stay running.