LXC, Docker, and the future of software delivery (LinuxCon)
slideshare.net
slideshare.net
The one phrase that's correct is "it's chroot on steroids". That is in fact exactly what it is. The exception is, it's even less portable than just a chroot environment. Docker adds extra features on top of the chroot, but that's basically its core functionality.
So the first thing you have to ask yourself is: does my software need to be run in a chroot environment or a VM isolated from all other applications? If no, you very well may not need this at all for your software deployment. If anything, Docker images create a bigger burden on your deployment as you have these large images to distribute, modify, manage. Of course they built in some fancy network transmission magic to make it only copy changed parts of an image, but this is still wildly less efficient than traditional means, and you still have to fuck around with the image to make it incorporate your changes before you push it.
If the big selling point is "commoditization", keep in mind that basically everyone rolls their own environment and customizes their deployment. It's the natural order of having your own architecture that fits your application. The one thing that's never going to happen, is you taking Docker images from the internet and never modifying them. This universal container system goes out the window the first minute you have to start modifying everything to fit edge cases, which is always going to happen.
This is why the Dockerfile is so important. I would never use an opaque binary container as part of my production app. But I definitely would use said container if I could point to its exact source repository, modify it to my liking, and rebuild it in a repeatable way. That is possible with the docker build system :)
You will no doubt point me to some script you wrote as a sysadmin. I wrote similar scripts as a sysadmin. They are unusable to anyone outside your immediate team of sysadmins - you can't use my scripts, and I can't use yours. That's because they're ambiguous. They assume that other things are already present on the system, and don't specify how to make sure that's true. There's a README next to them explaining what they require, how to bootstrap a build and so on.
Dockerfiles don't require a readme. They are not ambiguous. That's why they're useful. They also specify where to drop source files in a container, which allows you to get rid of your custom scp + flip-a-symlink ghetto deployment script.
Please don't pretend that a Dockerfile is some new paradigm shift in how to deploy an app on a remote host. And really, who gives a shit about "custom" vs "standardized" argument passing? And what's with the ghetto references? You do realize under the hood your Dockerfile and my 'ghetto script' are calling the same syscalls, making the same network connections, doing the same i/o, right?
Who gives a shit if my 'script' is 'ghetto' if it's ten times more user accessible, portable and backwards compatible than your proprietary format? A developer isn't going to spend more than 5 minutes looking at what the assholes in ops have given them to do their job; they learn it, they use it, they move on with writing and testing their code. Docker doesn't improve anything or provide anything you couldn't already do before. It just creates a new niche.
Your whole argument for the use of Docker seems to be "we are superior, because our operations are fancier, and we have created a gold standard." Face it: it doesn't matter how the sausage gets made. Docker just prescribes one way for how the meat is ground, vs all the others that work just as well.
Of course, you can achieve something similar with shell scripts, e.g. "if Java is not installed, install it, otherwise go to the next step". But if you get it wrong, or if there is some side-effect, the end result (after running the script multiple times) won't be exactly the same. The Dockerfile (with snapshots at each step) ensures consistency.
It's not a big deal if you are a shell expert and write that kind of script for breakfast, but it's not the case for everyone :)
Compared to what, though? Of course if your environment is homogenous enough that you can express your application as, say a Jar or a gem - then you are part of the lucky few and may indeed not need docker, because you share enough context with your target infrastructure that a lot of the bits are already implicitly deployed. (In other words: someone else had to move a big-ass system image around so that you don't have to).
But the typical application stack is not like that. It is heterogeneous and custom, and the only practical way to ship it reliably is to ship the entire system with it, because you've run your tests on a particular libc, postgres and ruby, built by a particular gcc, etc. In that case, your options are limited: 1) ship a VM or 2) ship system packages.
And if your current options are indeed to either ship a VM or system packages - then Docker suddenly doesn't seem that heavyweight after all :)
The Docker image may include several of those subsystems, or cross multiple of them, and so it now needs to be flexible enough to change one or several of those parts before it can be pushed to production to fix a bug. And what production environment will it be applied to? And is it possible that indeed you will have several images that are almost the same, except for key parts that can't be easily handled by yet-another overlay? (How many overlays will you have? Will you eventually stop adding overlays and redo the image to include the fix? What else will that affect?)
Complex systems require complex interaction, and Docker images do not allow for that; they are monolithic, all-or-nothing changes which can only be "modified" by either remaking the entire image (expensive), or adding another layer of overhead on top, which I don't think anyone has ever investigated to find bottlenecks or overhead problems.
To put it in simpler terms: Cfengine delivering a single change on a single file to a dynamically-assigned set of nodes is a lot faster, lower overhead, and direct than deploying a Docker change.
Cfengine delivering a single change on a single file to a dynamically-assigned set of nodes is a lot faster, lower overhead, and direct than deploying a Docker change.
To container-based architectures' credit, cfengine-style changes are not as 'known', since there are aspects of the environments that may differ. Containers can be 100% accurately cloned so significant testing can be performed before deployment. Of course, you could test with cfengine as well, it's just that when you attempt to set up nodes to do the tests against you'll suddenly realize the utility of container-based virt ;)
This has been the catch-22 i've had problems explaining, so bear with me here. Your docker container and my cfengine change are exactly the same thing in terms of testing and guarantees of sameness.
So, you have a Docker, and you hand-built it to an exact specification, and made an image that does a specific thing. You ship that image to 1000 nodes and it works "exactly the same" across all of them.
So, I have Cfengine, and I hand-configured it to an exact specification of operations, and made a config file that does a specific thing. I ship that config to 1000 nodes and it works exactly the same across all of them.
Why did they both end up giving me exactly what I expected? Because I controlled the input and the structure of the files to exactly my specifications, and nothing changed in those files once it was copied to the nodes, so they worked exactly as I expected. We can also assume I had no other dependencies for either to work.
(The Cfengine dependencies were the cfengine binary, which we can ensure is the same version across all hosts pretty easily. The Docker version, and the whole "ecosystem" of docker tools, is similar to enforce, though the much wider range of dependencies will make it much harder to support the same ones on all platforms, which is where we have to start using a unified platform, which is where the whole "run it anywhere on anything" bullshit becomes painfully unrealistic)
Uh-oh, here comes trouble! A developer comes over and wants 10 of the 1000 nodes to work a little differently. OK, we make a config file that Docker will read to configure the service a little different. Separately, we can make Cfengine read a similar config file that tells it how to configure the node differently. We ship the config file off to the 10 nodes, and Docker/Cfengine reads it in and changes its operations.
All of a sudden we have unexpected behavior due to the new instructions in the config file. One hundredth of our hosts are doing something different and we have to now account for that in anything else that happens in the future. (That is not some drawback of some technology, btw; that is called "I have a real network where things work differently across it").
The practical differences between these two is that the Cfengine change can be tested and applied quicker than a Docker, it doesn't force you to choose how to update your application ("am I updating the image? am I adding an image? am I changing the Dockerfile? am I changing the config? should I restart the docker or the service? do I have to reconfigure the autodetecting whatsit?" etc), and instead of hiding changes in a "Dockerfile" which is only parsed at build time, it keeps changes forefront in a configuration parsed at run time. You can make either one work just as reliably as the other. But one is a hell of a lot clunkier to get working in all cases than the other.
Also, i'm comparing a configuration management tool to a chroot environment network transfer tool, so take all of this with a grain of salt.
Uh-oh, here comes trouble! A developer comes over and wants 10 of the 1000 nodes to work a little differently.
In a formalistic and most desirable case, the prescribed node changes should be specified, versioned, documented and pushed. By the developer. With no input from systems administrators or operations people. The best way to do this is usually to define it holistically. Your approach emphasizes the speed of cowboy modification of running machines, whereas a formalistic (and continuous deployment ready) approach emphasizes the security, identifiability and repeatability of instituting the same environment through a standardized process. Both have their benefits, but if the latter is automated it shouldn't be any slower and would still offer benefits over the primarily manually driven PFCT-based ad-hoc change style approach you seem to be championing. Hate to put my finger on it, but an example would be that your salary and/or availability is no longer required.
All of a sudden we have unexpected behavior due to the new instructions in the config file. One hundredth of our hosts are doing something different and we have to now account for that in anything else that happens in the future.
That's just the point. By having a known entity that is versioned and self contained you don't have these issues. By making manual, ad-hoc changes with a post-facto configuration tinkerer (PFCT) like cfengine, you are out on a limb.
The practical differences between these two is that the Cfengine change can be tested and applied quicker than a Docker
Containers are very fast. I don't think you can make this argument without citation to numbers, nor do I think it's often relevant as speed of deployment, within certain boundaries, is rarely a bottleneck.
it doesn't force you to choose how to update your application ("am I updating the image? am I adding an image? am I changing the Dockerfile? am I changing the config? should I restart the docker or the service? do I have to reconfigure the autodetecting whatsit?" etc)
A formalistic release process that incorporate unit testing and the provision of test infrastructure, particularly that for systems including multiple, versioned services that must operate in tandem, requires some overhead. This overhead is about enforcing segregation between developer output and (automated) operations concerns to support continuous integration and continuous deployment.
I'd suggest that right now we're going to see this same sort of issue turn up regardless of how we choose to manage services (except manually, in an ad-hoc, error-prone, oldschool, cowboy style fashion), and in future we're going to see less and less systems administrators performing manual processes because some automated segregation mechanism or other will become popular enough to reduce operational demand in many companies for that skillset.
Quite separately to the above, I agree that docker fails to provide much of a meaningful feature set in this area, and that docker files might not be ideal. However, something will come along and fill in the gaps, eventually. Probably incorporating corosync/pacemaker or other, proven, prescriptive, cluster engines with self-healing / high availability features that rely upon continuous deployment of well segmented and versioned service instances for their very feasibility.
Unless you compose a complex system by composing multiple containers - which is the whole point of containers.
> they are monolithic, all-or-nothing changes
Compiling a binary is also a monolithic, all-or-nothing change. If you dig deep enough there is always such a change. In the operation of distributed systems (which you claim to be an expert of) that is a desirable property.
> which can only be "modified" by either remaking the entire image (expensive)"
Docker caches build steps. Which means it only rebuilds the layers which need to be rebuilt.
Typically this means the application code is rebuilt when the developer pushes a new version, while the underlying layers are left untouched. You know... the same "rebuilding" your ghetto deployment script currently does.
> or adding another layer of overhead on top, which I don't think anyone has ever investigated to find bottlenecks or overhead problems.
We have been using aufs layers as the build mechanism for lxc containers at dotCloud for roughly 3 years. In that time we've probably deployed half a million containers, and served a few hundred million uniques (and I'm being conservative). These containers included app servers, databases, and everything in between.
So, yeah, they've been "investigated" for bottlenecks and overhead problems.
2. Compiling a binary is not a 'monolithic change' (?), it is a static configuration with dynamic elements. And please quote the line where I called myself an expert of anything.
3. Uh, I don't have a "ghetto deployment script", but thank you for the kind words. Like I was saying before, your only option is to continue adding on more unions every time you change a file, which I bet will lead to performance degradation, if not just general application headaches in the future. If somehow Docker also prevents the need to stop and restart the container when adding a new layer (which should be possible with Union), that's great too! It still leaves a world of cruft behind in the form of old filesystem layers and a maintenance hassle.
Awesome, you have real world performance numbers! So how many layers can you add to a container without it going down before there's performance degredation? How does it affect memory, or disk space from added layers? Does aufs have an upper bound on the number of layers, or any other metrics? Would love to see some numbers on this.
Slide #44: Docker roadmap towards 1.0 seems to dodge the question of significant differences in function with regards the apparent plan to adopt a variety of storage backends with different capabilities, use of different virtualization environments as targets, etc.
I support docker as a project but I still really think you guys need to stop and ponder your architecture and goals before charging along too far. For projects to survive long term and be useful sometimes separating concerns is necessary, and I would suggest that's perhaps not being done well at present with some one-size-fits-all assumptions that are pretty anti unix philosophy (do one thing and do it well). What is the one thing? Is that really a general need? In all cases? What does a user lose with this abstraction? Rather than increasing scope, what would happen if you tried lopping those bits off entirely?