Why is Docker the new craze in virtualization and cloud computing?
opensource.com
opensource.com
This is it exactly. When I can develop locally using my favorite editor and my favorite OS, but then deploy -- still locally, mind -- to what is more or less the production environment... This has huge benefits for avoiding the "works on my box" problems that arise as code makes its way to prod.
In the long run I hope we start seeing distros or tools targeted specifically at building minimal Docker images. E.g. if I'm going to deploy a Ruby based web app in Docker, the stuff I need present in the container is _very_ limited.
Note docker isn't the only answer for this, for instance, you can use ordinary VMs with vagrant, but docker does it at ludicrous speed.
Also, production data has to be copied over separately?
Are these correct assumptions?
PS: I'm still trying to understand how docket/containers may fit in current process.
There would still be some differences, like if the kernel version on your local machine didn't match.
Note: I haven't tried this. Maybe it's a bad idea.
Sounds a lot like FreeBSD's release process, though personally I find that limiting!
The problem is of course that 'very close to' is not 'equal to'.
My own thoughts in this space are detailed at http://stani.sh/walter/cims/ and http://stani.sh/walter/pfcts/
I've been using vagrant to fix exactly the problem of "works on my box" for years. Docker is way too much to explain to my small development team. They struggled to deal with vagrant for some time before figuring it out.
Vagrant is perfect for us right now, and it would take a force of nature for me to even begin to push docker on them.
The ability to version your images and easily build on them is nice too.
First time provisioning of a box can take up to 30 minutes even with a non-mature app. Mature apps with many moving parts can easily get near an hour- using Docker, you can bring that down to 5 minutes or less, with the majority of that being downloading the containers.
Neat thing though, is that you can use Docker WITH Vagrant to do this. The setup I'm doing on my side projects, essentially goes like this:
If I'm developing, I use Vagrant to spin up Docker containers(using Docker as the provisioner) for supporting items- databases, memcached, things like elasticsearch or a message queue- then develop from my machine(a mac). If I want to get someone else in on the project, I don't have to care one iota about what they're on, Docker will make it work.
On top of that, there's the capability to have your app itself running in Docker containers for development. This means no having to worry about dependencies, environment variables, etc- and if you keep things up to date, you can even just upload your app to a private Docker repo, and use that to deploy to your server(though in a team this may not be the most realistic option at the moment, I've not given it a shot, though I don't see why it wouldn't work).
For me, Vagrant is the tool I use to tie everything together for development. Docker is what I have it run off of, instead of Chef/Puppet/Salt/Ansible, so that my build times go from 30 minutes to whatever download speed I get.
Of course then you risk developers of libraries/tools only developing/testing against that particular environment, but that could be solved by having a variety of Docker/Vagrantfiles integrated with a CI system.
If used properly Docker enforces pushing those upgrades back up the pipeline as far as they should.
I see a process where ops are responsible for the base images (base os images, and layers for different software stacks) and can handle security updates at will. Developers can maintain small dockerfiles that just install their application on top of the ops-managed image. ops can then be in charge of building (and re-building) the images and deploying them. Of course, that can all be fine-tuned to your liking.
This is one of the best explanations of Docker I've read. (Altho I suppose it requires knowing what Hypervisor is, if just at a topical level...perhaps the same for a "kernel")
Docker is best described as application virtualization.
Docker is now a wrapper around libcontainer which runs an executable as PID1, with a build system for those containers (Dockerfiles) and support for features exposed by libcontainer (ports, bind mounts, etc).
The use case of "operating system level virtualization" is "I'm already running Linux, Solaris, BSD, AIX, or whatever, but I want to run another copy on top of it in order to segment off my clients/webserver/business unit/whatever". There's an implication that there'll be a real init as PID1, you may want to run sshd or another access point to let users in, and it's a "pet" in the pets v. cattle parlance.
The use case of Docker is "I built my application with this set of libraries on top of Ubuntu, but I want to deploy it alongside applications with different versions of libraries (also on Ubuntu or RHEL or whatever), and my application is the only thing that'll be running".
Remember when you got tarballs to unpack in /opt/${application} which came with start scripts which started with "export LD_LIBRARY_PATH='/opt/oracle/lib'"? Docker does that in a modern way. It can also do more than that with etcd, fleet, haproxy, and bits tacked on top to shuffle containers around, but that's the core of docker.
You, the developer, ship your application along with all the libraries it needs and it runs in a container on top of Ubuntu but it's using RHEL glibc/libwhatever without starting logind and all the other services associated with RHEL.
But the libcontainer networking stuff and integration with cgroups does provide more segmentation than chroots, and the networking parts are nice.
Granted, I think the best possible use case for Docker is in shipping fat apps (ala OSX or Windows) so Spotify for Linux can run on anything that supports Docker instead of anything which supports dpkg, but eh.
I totally understand the enthusiasm around OS containers. I forget sometimes that this is a new thing on Linux.
Running Solaris and FreeBSD is like living in the future!
[1] The full quote ends "...poorly", but Docker seems to done well-enough, just rather late to the party.
The functionality has existed (at least primitively) on Linux for some time, but culturally Linux admins have been more drawn to HW level virtualization instead.
OS-level virtualization has been a part of the FreeBSD and Solaris cultures for much longer.
Another fair argument is that Docker does such a good job at abstracting the configuration and management of OS-level virtualization that it truly changes things.
Maybe. But if that's true, it means that Linux admins have ignored this hugely useful technology because it was hidden behind an impenetrable wall of text-based configuration.
If two servers require different versions of a library, I see that as not a use case but a problem of technical debt, to be solved by repackaging and retesting with the correct version.
> export LD_LIBRARY_PATH='/opt/oracle/lib'
I guess some people are required to run such badly-maintained software, but I don't see why anyone's excited about the prospect.
It makes sense for me, as a developer, not to go through repackaging and retesting for 7 distros with 4 different packaging paradigms and 3 init systems when I could simply tell you to run a container. None.
Disk space and memory are now cheap. Statically linked binaries are coming back. Containers as the new /opt strike a middle ground between distro portability and developer effort. It must be nice for you to live in an ivory tower.
And if I had a goal of minimizing any interaction between my code and its environment, Linux syscalls are a much bigger API than I would choose. It's not as if you can claim to support a distro whose kernel or docker version you haven't tested on.
Think how ridiculous it is that in many cases today we run an entire simulated computer, with a full general-purpose operating system, just to power a single application (e.g. web server, database, etc).
If you're using Docker in development and then not using it in production, I hope you're using chef too, or something like it. The point to me is embracing the whole idea of "replaceable, reusable, disposable" components and DRY. If you're repeating by hand any complex steps in production that you also did in your development environment; never mind the time wasted, you now have twice as many chances to make mistakes.
http://serverfault.com/questions/261974/how-much-overhead-do...
These are good examples of there is no such thing as "the" virtualization and even front runners have a near 10:1 ratio of performance.
Its like talking about "the" sorting algorithm.
Its especially tough since there is no "the" benchmark. So if you're running a compile farm two years ago, the first article is quite relevant, otherwise maybe not.
I've done a lot personally and professionally with virtualization and as a gut reaction you can usually justify and afford better specs because not all images will max out CPU at the same time, etc etc so in terms of production "stuff" done per hardware dollar you get a little more done with visualization assuming a sensible architecture. Then again you blow more money and labor and labor is money, on software for virtualization. Then again management and backup is usually simpler with virtualization, except when it isn't like when vmware and the NAS kills your mysql instances on an image when you run a backup and it starves the kernel for 60+ seconds of IO whoops.
On average it ends up being about the same performance as non virtualized for a given amount of $$$$ but its much more flexible and responsive to change, which sometimes is worth it.
So, let's say I have a fairly typical setup, 1 web app, 1 database. Normally I would have 2 VM instances to start with. When I need to scale I add more nodes for web and/or db... both are load balanced, of course.
How would docker help me? I guess I should google for docker use cases, but I thought I'll try my luck here too.
Test/dev environments are pretty close to production ones.
It seems there is a lot of hype, but I fail to see it through. I'm wondering if all this hype is artificial, time will tell of course.
Thanks.
On the one hand, containers are more efficient than virtual machines. Hence, you get more efficient usage of your hardware. Also, you can pack more containers in a given host than if you use virtualization. Net result: less hardware to run the same services. In turn, this also readuces the network load because the more services in the same machine, the less traffic on your network and better latencies between colocated services.
On the other hand, containers (as envisioned by docker) are much faster to deploy. You build images which are compositions of layers. You have a base layer containing the base OS, and then a "stack" layer containing your application stack, and finally your app on another layer. An image is a consolidated view of all these layers, but the layers can be fetched and managed independently. Hence, you can deploy containers much faster (no need to re-download every layer each time you want to deploy a new service). Additionally, they build faster and they launch faster (nearly as fast as a non-containerized application). This allows you to do even cooler things such as socket-activated containers, which would be painful with the startup times of virtual machines.
Obviously, this blends very well with the idea of building applications in terms of distributed services (or components or whatever). As a result, there are multiple projects pushing for "cluster-aware" (replicated, varying levels of consistency) services. People from the virtualization camp was already pushing for that, but now it is even more important because containers are leaner and cheaper to deploy.
In my mind scale means use more resources, but in this case I fail to see benefits, except quicker env setup for the app, but for that there are automation services.
Yes, cluster-aware apps, if I want to scale 1 app, I'll create dedicated nodes/vm instances for it, how are containers helping? Why woudl I need multiple containers on the same host, for each app instance?
Thanks again.
Why would you need more containers on a single host? Because apps are usually not uniform. If you dedicate one host per service, then you need hosts to serve the maximum computation power at the peak consumption required at any point in the app's lifetime. Also, VMs are heavyweight, which means that it is painful to break your app into small, independent services that you can spin up and down quickly. Containers are much better at enabling you to modularize (and thus be able to scale on a finer grain) your application.
Now you may argue that you can spin more AWS instances when under load. But spinning up VMs takes time, much more time than spinning up containers (we're talking on the order of minutes vs milliseconds here). This may seem a small difference, but pair it with socket activation for instance and suddenly you go from having to carefully manage your scaling to getting it quasi-automated for free.
Finally, there's the ecosystem. You can create dedicate nodes/vm instances, sure. But watch a demo of CoreOS's fleet automatically and nearly-instantly spinning up and maintaining coordinated containers [1] and you may get a feeling of why people is so excited about this whole thing.
[1] http://coreos.com/blog/cluster-level-container-orchestration...
My current pain point is the hdd space, since I have to resize VMs storage when needed. Even though it was automated, but still I'd prefer to use some shared hdd space and not to worry about it.
Interestingly most "real virtualization" requires somewhat recent hardware but LXC containers don't take much if any, so "free machine off junk heap" is good enough for experimentation.
Historically there have been strange intercontainer isolation problems so don't assume its as perfect as hardware assisted virtualization.
Also you're sharing a kernel which is both good and bad, if you were hoping to test out a new kernel or run an entirely different OS, thats too bad.
Edited to add, another fun analogy if you like the chroot analogy, is when spinning up a process was glacial on windows (but not too bad on linux) that lead to the development and push to use threads which were a huge win on slow windows not so much on linux but we got dragged along anyway. In a similar manner, full virtualization is really slow to deploy and spin up and spin down, containers are smaller and weaker but much like threads vs processes are much faster to start up or shutdown. A really low latency or fast spinup spindown virtualization tech would probably wipe most demand for containerization. Or what I'm getting at is, spawning processes on windows was super slow, so we got threads, and spawning full hardware virt is slow, so we get containerization aka chroot++