But it's the less common choice.
I am guessing convenience is more important than the better solution that would ultimately be more just as convenient and more efficient if it gets enough eyeballs?
But it's the less common choice.
I am guessing convenience is more important than the better solution that would ultimately be more just as convenient and more efficient if it gets enough eyeballs?
At this rate we'll end up shipping containers as 'apps' to the clients machines with a suitable emulator at some point.
All this luxury comes at (considerable) cost and not everybody seems to be doing the math before deployment which more often than not leads to terrible efficiency.
But we're technology fans, so 'oohh! shiny!'.
While VMs and containers in VMs don't utilize the hardware as efficiently as a dedicated server they make your people much more efficient and happy. Unless you are running at a fairly large scale making your people more efficient gives you a much better ROI than making your servers more efficient.
[1] https://www.joyent.com/developers/videos/docker-and-the-futu...
Every dedicated server that you set up has an opportunity to not be duplicated perfectly. Every system that knows about your dedicated server has an opportunity to hard code what it shouldn't. Which makes these a potential point of failure. Add enough of those, and you're statistically guaranteed that the careful architecture that you have for failover is a pipe dream.
If everything is deployed with containers and discovery and the correct provisioning, then dealing with the fact that containers move around forces you to solve all of your other problems. And containers provide an abstraction layer that makes the rest of it straightforward.
Let me illustrate with an example.
When I worked at Google in 2010, I remember reading an article from eBay about how they finally manage to transition everything off of a running data center, without interrupting live traffic, and how much planning it took them. And they were congratulating themselves on what a heroic feat they had managed. Most companies today would still consider that a pretty amazing feat, and would find that challenging.
At the same point time I was learning how things were set up at Google so that you could drop any data center at random with barely any interruption of live traffic, and with no manual intervention required. And Google occasionally does this without warning to important data centers just to be sure that it works.
I like the test-the-panic-button attitude that google brings to these things, I've yet to get someone to accept my challenge to power down their supposedly automated fail-over solution, it's supposed to work but they usually can't be sure or management would surely not allow such a rash thing as a live test and it's bad form for me to then walk up to the switchboard and trip the breaker. Verry tempting...
This is, literally, the reason.
You can replace rails with any similarly bad technology. I got this explanation a few weeks ago at my job:
the java build process (I'm not kidding) has such a complex dependency graph that it must spin up full containers to do each build.
----
If dynamic languages supported better packaging/isolation this entire thing would be research projects.
(I am reading through the comments trying to work out if I should learn Docker. I already know how to use a virtualenv and I already know how to use a VM.)
One of the preferred ways to deploy rails is (was? I did this ~2 years ago) was to check it in with git and the name/version of the ruby environment to use is stored in the file ".rbenv-version", with dependencies managed by bundler (which you point to a local gem server). Install is then 1) use rbenv/ruby-install to install basic ruby, 2) install app with "git clone", and 3) run bundler. Many tools exist to do this in one step over ssh/etc automagically.
Even better than rbenv, you can just use chruby[3] to point to any rubies you want; just check one into your project itself (or whatever) and configure the siteruby/etc load paths to point to project directories. Really, chruby just fixes up your dev environment to point to a specific ruby; you set the actual project to be self contained with known paths, just like you would do de facto in a container.
While dependency issues were a problem back in the ruby 1.8 / rails 2.x days, this
[1] http://bundler.io/ [2] https://github.com/sstephenson/rbenv [3] https://github.com/postmodern/chruby
At least with the Borg containers (I'm not familiar with how Omega and Kubernetes do things), there wasn't any additional layer of abstraction - the fact that there were multiple jobs running on the same kernel wasn't hidden from those jobs (although they didn't have to be aware of it). The containers were purely used for in-kernel resource accounting and control.
Assuming you're operating at scale I don't see why that would be the case. And if you're not, what's the point?
> The more different things you can pack on a machine while still ensuring that the high-priority/low-latency jobs get prompt access to the resources that they've reserved, the higher overall utilization you can achieve (and hence bring costs down),
Yes, that's the theory. But in practice you're assuming better static control over the situation than the operating system running multiple jobs will have over the dynamic situation. So you'll need to over-provision and then you're back to square one with your utilization or alternatively you'll under-provision and then you will run into performance issues. TANSTAAFL.
(For instance, what's to stop each container to ship another implementation of the same library as a dependency, say SSL).
A lot of user-facing services at Google have to be over-provisioned in order to handle the cyclical usage patterns (the daily query peak is far higher than the average for most services) and to be able to survive the loss of a datacenter or two. This results in a lot of under-utilized servers for a big fraction of the time. So by packing lots of medium and low priority jobs on those same servers (and over-committing the resources on the server), you can soak up the slack resources; in the event that the resources are needed by the user-facing service the kernel containers ensure that the all the less latency-sensitive jobs on the machine don't compete for resources with the user-facing services.
It's true that the performance isolation when there are tens of jobs running on the same machine isn't going to completely match the performance isolation of running a service on a dedicated server, even with kernel resource isolation via containers, but you have to make cost trade-offs somewhere. The number of Borg services that could justify requesting dedicated machines was very small.
And to address your other concern about the OS not having so much insight into what's going on - Borg containers consisted generally of a single process, running on the machine's normal kernel. The containerization was just for in-kernel resource accounting/isolation. (Using Linux control groups, rather than anything fancier like LXC or Xen)
Thank you for the insight into the number of processes inside a typical Borg container, so that was basically a kind of 'heavy process' rather than a complete application with all dependencies (including other processes the main one depended on) packaged in, this is something I wasn't expecting at all.
Really?
My impression was that typically you'd have a process for the service, a borgmon process for monitoring, and maybe another process to ship logs off in the background.
Developers would only think about the service process (which itself typically was a fairly thin shim in front of other services), but a borg container would have more than that going on in it.
The logsaver would also be a separate job, although typically running co-located 1:1 with instances of the actual service job. The service and the logsaver would have access to the same chunk of disk (where the logs were generated) but otherwise they were separate as far as the kernel was concerned. (As far as Borg was concerned they were very much related, but that was at a much higher level than the kernel).
In another view, it's another approach to what many look to Chef, Ansible, and Puppet to do. Combined with something like Mesos or Kubernetes, you can quickly deploy to a heterogenous cluster, without a lot of install scripts running.
Some of the other uses cases, such as running multiple containers simultaneously on the same hardware, make less to me.
Here's my answer for "why": DRY. Once you've deployed hundreds of servers using the same exact Ubuntu 12.04 LTS kernel base, why not just completely abstract the OS away and focus the attention on scaling the OS services that matter? Why is that when I decide that I need to scale out, I need to copy every library of the OS and every line of code for the kernel and redeploy it every time I add a node?
> At this rate we'll end up shipping containers as 'apps' to the clients machines with a suitable emulator at some point.
That's exactly the point. Care to elaborate on the downside of such a promise?
Because it adds a layer that makes no sense unless you have very specific use cases. Though I see the point regarding people efficiency, that one makes good sense (see other comment in this sub-thread)
> Why is that when I decide that I need to scale out, I need to copy every library of the OS and every line of code for the kernel and redeploy it every time I add a node?
If you're doing it that way then you are simply doing it wrong. See: chef, configuration management and various deployment services (of which you could argue containers are one off-shoot, but they focus (imo) on the wrong level for all but the largest companies). Containers are like sandboxes with significant overhead for applications that focus on ease of deployment (but that's strange to me because I see that as a one-time cost for most of my own use cases, though I can see how that equation would change if you deploy lots of things configured by lots of different people to a single set of servers, especially if there are conflicting requirements between those deployments).
> Care to elaborate on the downside of such a promise?
That's my personal view of hell, if you don't see any downside there please ignore my vision and continue as if nothing was said.
From open, text based standards to shipping arbitrary binaries in a couple of decades. And I thought GKS was about as bad as it got ;)
A chef script is basically the automation of "I need to copy every library of the OS and every line of code for the kernel and redeploy it every time I add a node?" I'm sorry if you didn't pick up on my implied remark. Two problems are then introduced when automating those actions: (1) it doesn't negate the fact that I need to store and deploy a 700M sized OS layer every time I want to add a node (which takes minutes, not seconds with non-containerized configs) and (2) maintaining config scripts can (not always) be painful (version control, rollbacks, etc)
> unless you have very specific use cases.
> Containers are like sandboxes with significant overhead for applications
Again, do you have experience using containers? You seem awfully dismissive ("you're doing it wrong!") in a way that suggests that you might not entirely understand how they actually work...
In a nutshell: running a 'standard' combo of apache and a DB server as well as some auxiliary bits and pieces inside 'containers' a year ago gave significant overhead compared to running those without the containers. I'll re-do this and I'll probably do a write-up because the subject is interesting. This comment and follow up (https://news.ycombinator.com/item?id=9567623) are by people using this tech in production right now and their experience echos mine (but they're very far down the line compared to where I stopped).
Besides that particular use case (where performance and isolation are the key components to be looked at) some interesting points have been made in this thread which has shifted my stance on container use depending on what the situation is. So I don't think it is valid to classify me as 'awfully dismissive'.
FWIW I have not used containers in production (yet) but I'll be more than happy to if I can figure out where and how they can bring me an advantage, which is pretty much how I approach all tools.
That gives me both environment separation (A needs Ruby 1.9, B needs Ruby 2.0), resource accounting on a per-app basis, and a repeatable foundation in case I need to re-deploy the server or spin up new instances.
Disk - the technologies used for disk isolation (save chroots) are very poor performance, and in some cases can cause resource contention between what would otherwise appear to be unrelated containers. As an exmaple, using AUFS with Node creates a situation where any containers running on the same file system can only run one at a time, regardless of the number of cores. It's silly. Device mapper, on the other hand, is just plain slow (and buggy, when used on Ubuntu 14.4).
Network: The extra virtual interfaces, natting, and isolation all come with a performance penalty. For small payloads, this manifests as a few milliseconds of extra latency. For transferring large files, it can result in up to half of your throughput lost. Worse, if you have two docker containers side by side but due to your discovery mechanisms one container uses the host device to talk to the other container, you create what is known as assymetric TCP, which can cut your performance by a fifth or more. Try it out sometime, it's entertainingly frustrating to figure out.
Security: My favorite. What's the point of creating a container for your application if you're going to include the entire OS (and typically not even bother to update it with security patches). A real simple DOS on docker boxes would be to get the process to fill the "virtual" disk with cruft. You'll impact all running processes, the underlying OS (/var/lib/ is typically on the same device as /), and create such a singularly large file that it's usually easier to drop the entire thing and re-pull images instead of trying to trim it down.
Sorry if I sound down on the tech, but I've been fighting to make this work for production, and all of these little niggles are driving me batty.
Docker is fun and great when it's running on your workstation and coddled by your fingers at the terminal, but there's a lot of gotchas and missing parts when it comes to putting things into production, to be taken care of in a hands-off manner. There still isn't an easy way to centralise logs from a container app's STDOUT. Yes, there are other containers you can install to ship logs (which work for the author's use-case, not necessarily yours) or you can hack together something horrible. If you want to look at container logs, you have to have root rights. You can be in the docker group and have full control over the daemon, but the container log location is root only, and is made afresh with every container. (and don't forget to rotate those logs!)
My latest fun with docker is that one of my docker servers, built from the same source image and running on the same configuration plan in ansible as my other docker servers, fails to start docker on boot. Some sort of race condition, I assume. Basically it fails to apply its iptables rules and dies. People talk about making problems go away with docker, but it's a trope in my team that any day I'm working with docker, I'll be spamming chat with problems I'm finding in it from an ops point of view. And I'm just a midrange sysadmin :) But the point is that adding Docker adds an extra layer of debugging. The app stack still needs to be debugged, and now there's an extra abstraction layer that needs debugging.
Plus, in my particular case, there's the irony of using single-function VMs to run a docker container, which is running the same OS version as the VM :) (my devs bought into docker before I arrived...)
I can see some (mostly potential at this point) security advantages but that's about it (and maybe those advantages will be enough to justify the performance overhead but containers are mostly treated as a silver bullet by the adherents and I'd like to see a bit more balance).
No. A container is just a tarball of user-space code run with some isolation. The kernel is still the kernel. Run multiple containers on a machine, and the OS manages all of their processes at once.
And I just verified, you can kill a running process from outside a docker container. So the OS does see it and probably can do all its scheduling magic.
How does this perform in practice when they start talking to the outside world at or near capacity? How does it perform when they start talking to each other using some defined interface? (But presumably, no longer regular IPC).
So if one or more active containers could share resources then they won't, which leads to inefficiencies because you'll be running a much larger number of processes than you would otherwise (because of duplication) requiring a larger memory footprint and probably less efficient cache and/or IO utilization.
The deployment of the apps will be easier (which is a definite plus) but machine utilization will be lower and the amount of software running on a single machine will be far larger than otherwise, especially if multiple versions of dependencies are present on the same system.
A container is very much not a single process, it can contain many processes and some of those processes will likely duplicate components in other containers but without the resource optimizations that a kernel can normally perform.
Although true, that probably isn't really very significant compared to the vast wasted resources of idle dedicated machines. Which is hard to avoid without the vast wasted resources of a highly paid somebod(y|ies)
I also don't quite understand how one can reserve CPU cycles, memory and deliver IO guarantees without the same over-provisioning that you'd have to do using regular virtualization. After all, as soon as you make a guarantee nobody else can use that which is left over, so in that respect I see little difference between virtualizing the entire OS+app versus re-using the kernel (ok, that does save you the overhead of the kernel itself but that's not a huge difference unless you run a very large number of VMs on a single machine).
In the event that there ends up being no best-effort resources available on a machine for a significant period of time (because all the user-facing jobs are busy and using their guaranteed resources) Borg will shift the starving batch jobs to other machines that aren't so busy.
Where regular virtualization runs multiple kernels (which in turn will run whatever applications you assign to them) containers appear (to me, feel free to correct me) as a way to 'share a single kernel' across multiple applications dividing each into domains that are as isolated as possible with respect to CPU, memory, namespaces and IO (including network) provisioning and allowing multiple version of the same software to present at the time without interference.
The CPU, memory and IO provisioning can be thought of as a kind of 'virtualization light' and the namespaces partitioning should (in theory) help to make things a bit harder to mess up during deployment.
Leakage from one container to another will probably put a dent in any security advantages but should (again, theoretically) be a bit more robust than multiple processes on a single kernel with shared namespaces.
So I see them as a 'gain' for deployment but a definite detriment for performance because it appears to me we have all (or at least most) of the downsides of virtualization but of course you can expect both virtualization and containers to be used simultaneously in a single installation with predictable (messy) results.
I'm really curious if there is an objective way to measure the overhead of a setup of a bunch of applications on a single machine installed 'as usual' and the same setup using containers on that same machine. That would be a very interesting benchmark, especially when machine utilization in the container-less setup nears the saturation point for either CPU, memory or IO.
And that's assuming that it'd work exactly the way you're thinking.
I feel like the win over running VMs (which incur something like a 12% overhead compared to both Docker and running right on the machine for a single application), plus flexibility, plus ease of deployment is worthwhile. I mean, the current situation is running VM images anyway, right? This is a step in the right direction over that, even you must admit.
But you've made me curious enough that I'll do some benchmarks to see how virtualization compares to present day containers for practical use cases faced by mid-size and small companies, my fooling around with this about a year ago led to nothing but frustration, it's always a risk to argue from data older than a few months in a field moving this fast and more measurements are the preferred way to settle stuff like this anyway.
The linux kernel does not "lose track" of processes/libs inside containers, they are simply namespaced, like a more extensive chroot environment.
Of course it does make it easier to package and deploy applications (and to ensure their correct application) but to pretend that there is no cost associated with this is simply not true.
There is also ksmd that is useful with VMs, where memory is at a premium, though I'm not certain it is compatible with lxc yet.
But in practice (at Google-scale, anyway), that's dwarfed by the efficiency gains you can get by squeezing lots of things on to the same machine and increasing the overall utilization of the machine. Prior to adding kernel containers to Borg to allow proper resource isolation between the different jobs on a machine, the per-machine utilization was really embarrassingly low.
Another point to consider is that not all jobs are shaped the same as the machines - some jobs need more memory (so if you put them on a number of dedicated machines adding up to the total amount of memory needed, there will be lots of wasted CPU), and other jobs use a lot more CPU and less memory (so if you put them on a number of dedicated machines adding up to the total amount of CPU needed, there will be lots of wasted memory).
By breaking each job up into a greater number of smaller instances and bin-packing on to each machine, you could take advantage of the different resource shapes of different jobs to get better overall utilization.
No, you use containers despite the fact that your hardware utilization goes down (mainly because no shared pages between applications), because your huge sprawling environment is too hard to change with flag days.
Being able to strictly apportion resources between the different jobs on a machine (and decide who gets starved in the event that the scheduler has overcommitted the machine) means you can squeeze more out of a given server (by safely getting its utilization closer to 100%)
There are other definitions of the word 'container' that are closer to 'virtual machine' and include things like a disk image which is much harder to share, but that's not what's being discussed in the context of Borg. (Not sure about Kubernetes, that's after my time)
I like you main point however: I would like to know, given vistualization has X% overhead, what is X?
Perhaps because that's not what happens, the host runs one copy of linux, which namespaces the containers.
Also, possibly a clairification: Unless I'm misunderstanding what you mean by "running several linuxes on a linux machine", I believe you may be mistaken about the way containers work. Only one Linux is really running. And that is the Linux that the kernel comes from. The other stuff doesn't run unless you tell it to (so, no init, no daemons you don't specify, etcetera). Yeah the image size can be a little fat if you don't trim them down, but you can have a container nearly as small as your code is, if you statically compile. On the order of just a few bytes of overhead.
you no longer have an OS in the traditional sense.
you just dont get the debugging stuff (which is okay as long as you can choose)
yes, really!
it does reduce the attack surface/amount of things.