How does Docker affect energy consumption?
arxiv.org
arxiv.org
Containers allow us to condense workloads in a single OS runtime –while preserving isolation– where otherwise the same workloads would have spanned multiple machines or VMs, each with its overhead and slack (unused resources).
Example: consider you need to deploy not just a single instance of Wordpress, Redis, Postgres, etc., but a complex application consisting of many of those components.
You can either choose to (a) deploy each in a different machine [incurring in the overhead of OS + unused resources]; (b) in different VMs [incurring in the cost of OS, but being able to share a resource pool]; or (c) in containers, sharing the OS and incurring in the overhead of the container daemon.
I would love to see an article that takes these points into account.
I'm starting to believe that the increased density provided by containerization is a myth, in practice. Both because orchestration tools bring their own overhead (compare all the proxies and filesystem/network overlays with multiple virtualized kernels) but also because containerization goes hand-in-hand with microservices thus increasing the number of components (often times with no real reason other than to be hip).
If you're running 20-30 containers on a beefy VM, you're not really condensing anything. You're just moving from running hundreds of small VMs on a single server to running much less larger VMs.
You might be dismissing microservices too quickly. They do have overhead but so does any level of abstract; the benefit of them though is clear separation of responsibilities between services and residency(Swarms, clusters, etc). Both can be achieved with Vms but VMs weren't built with these goals in mind
I'm not sure the toaster analogy will gain mass acceptance.
When Docker can safely protect a Minecraft server in one container from a local DoS attack coming from a bot running in a sibling container, I'll reconsider using VMs. :P
Of course you are condensing. Those small VMs would have been running a full OS runtime each.
Assuming you relocate those workloads onto a single machine (1 process = 1 container), you no longer run 20-30 copies of an OS, just a single one.
And the container engine takes the place of the hypervisor.
That's something but probably not as much as you think since hypervisors can share identical pages (e.g. the Linux kernel) across guests and the base footprint for a server Linux install is not that high as a percentage of the private data most applications use. Unless you're running a ton of unnecessary services on those guests or have an application which uses almost no RAM you're talking about a fairly modest percentage savings even before you factor in all of the things you might be running for container management and other overhead on that side.
The other thing to remember is that this works both ways: containers are great for being able to upgrade one component independently but that means that e.g. you might have a dozen different versions of a common shared library because not all of your containers are using the same base image & version and with Docker your storage driver might actually force shared libraries to be duplicated across all processes anyway.
With ASLR, I'm now sure the gains are that substantial.
Or, to put that another way: the host memory for most modern hypervisors consists of a heap of "new" pages, and then a generational garbage collector that moves said pages, if still alive, into a content-addressible "old" store.
As such, if two VMs each have a process that
1. calls malloc() 1000 times to get 1000 1-page buffers randomly spaced through their memory, the mappings different for each VM; and then
2. uses a fixed PRNG seed to generate random data [but the same random data] to fill those pages;
then those two processes' pages will still get collapsed together for a 50% savings.
Everything else is mapped onto the single Linux kernel, and many pages are shared as a result; I think that even libs are able to be shared across VMs; so if you had 2 identical versions of glibc in 2 VMs, only 1 would be loaded and used.
1. "application containers" are effectively a single process [though that can fork more] with some kernel process-struct fields set to nonzero values, indicating that the kernel should present this process a different view of its environment.
2. "virtual machines" are the processor providing a separate virtualized view of the CPU, on which is then booted another virtualized kernel, which brings up with it virtualized OS services and eventually an app.
3. Between them, "OS containers" are a hybrid: they start up all the userland virtualized OS services that a VM does, but they do so on top of a kernel that's not actually a fresh, separate kernel; but instead a kernel that has been told (through setting tons of containerization process flags) to present to this group of processes a view of the world where this kernel looks like a fresh kernel in a newly-started VM.
"OS containers" are basically a raw optimization over VMs by asking one kernel to pretend to be multiple kernels, and to manage one pool of memory instead of having multiple pools of memory. Anything you can do with raw VMs, you should (in theory, given good inter-container isolation+quota logic) be able to do with OS containers as well.
containers bundled up libs and binaries in a single package, only for an "app" to come reliant on a zoo of containers doing one little part of the whole.
Makes one wonder if the stack is made of rabbits rather than turtles...
Ultimately there's no reason why containerization can't be pushed right down to the language level. Consider .NET's AppDomains, or even further, a capability-secure programming language which isolates at the object level with zero overhead over ordinary languages.
I started my reply by listing a plethora of reasons why this wouldn't work. (For one, this would work because it does work, right now in Erlang.) But they all came down to that it seems like you're missing some of the problems that containerization solves. You can wrap up almost any service -- regardless of language or versions or runtimes or how it interacts with the filesystem or what versions of libraries it depends on or what global configs it expects or anything -- and ship it as a self-contained normalized service that can run right alongside any other number of other self-contained normalized services that require their own global configs and libraries etc etc even if they're incompatible and whoever is deploying them doesn't even need to care.
Everyone has their own language and toolset that they're comfortable with and productive in. That will never change. Containerization abstracts over all of them and normalizes their deployment. You can never get that with a solution at the language level.
Eh? Erlang has no equivalent to cgroups/namespaces, or even an equivalent of non-UID-0 code execution on its VM. There is in fact no isolation mechanism in the Erlang VM; all code is "privileged." Untrusted multitenant code execution is a pipe-dream for now, unless you graft on another sandbox inside Erlang, ala CouchDB's V8 C-port [and more recently luerl] sandboxes.
(I've been very much considering contributing code for "non-privileged Erlang processes" and "Erlang process namespaces"—adding things like "namespace outboxes" that will crash their own virtual nodes rather than flood peers—but it's not there right now.)
But if you're deliberately running malicious code even cgroups/namespaces won't save you from some attacks. Timing and cache attacks can be done without breaking out of the jail.
I would say one member having their private keys stolen[1] is a "fatal shot".
Yes, you can get cross-app deployment conflicts (packaging, etc), and limits your cloud deployment options, but it is definitely another option and has lower overhead than any of the first three.
It also creates fragility. What happens if a process is buggy and rallies up to 100% CPU? It affects all others.
Although to solve these issues you could use cgroups and namespaces... Aaaand we're back to containers again.
I've had slowdowns in the order of 30% in terms of requests per second on some of my stuff inside Swarm as compared to running processes outside containers (arguably with four or five moving parts and entirely anecdotal, but still enough to give me pause), and I'd really like to understand how to shave this yak.
However, you might see docker-proxy eating a lot of CPU due to increased traffic. Or dockerd adding overhead to collect statistics, manage log buffers, etc.
My point is, pure containers using basic kernel features don't seem to add a lot of overhead. It's the sugar on top that sometimes is the problem, but that will vary depending on the container runtime, the app being contained, extra features that were enabled, etc.
I've seen CouchDB perform just fine in a container when using Docker's "bridge" network and drop to a halt if using Docker's "host" network. It seemed to dislike the situation and was doing way more syscalls than usual. Just to show the app might not like a certain environment.
Even with some overhead, it's a trade-off I'm willing to accept considering the advantages in the development workflow and managing the infrastructure. Most of what I've seen are rough edges that will get ironed out with time.
(I've also investigated the Ubuntu fan technique to lower overhead, but am looking for a more definitive solution)
You can say that again. Tried running Docker container with large range of exposed ports once (because SIP)... Each port got its own docker-proxy and the system basically suffocated. Not a nice view. :)
CPU instructions are executed as-is. Running on VmWare, ESX, Xen, Docker, LXD makes no difference.
It's a different story with network and storage access. Impact of containerization/virtualization is variable(5 to 95% performance drop).
The main overhead is in IO though. Closer to 10% with KVM virtio than 95%.
Another thing to note is that redis and postgres are i/o heavy and many production environments (esp in aws) will choose not to dockerise i/o heavy stuff.
I wish they had included a statement like that so I didn't have to stick my neck out and give an eyeball estimate, but I suppose this kind of statement is harder to defend and the data they provided was more nuanced.
With that said, power is largely divorced from cost in actual operation except at extreme scale.
This is because purchaseable and billable units of compute are usually not utilised to 100% capacity, both in cloud and bare-metal situations. Another way to say this is, most people essentially prepay more "power budget" than they actually use.
Since a primary use of docker (esp via kube) is workload consolidation, it's hard to know what real impact this has on the world.
Generally containers vs VMs is a wash... except when considering security.
There used to be some issues with power saving in Xen that increased power use when idle, but have been fixed around 4.4 or so.
The comparison was done on
1) Ubuntu server 16.04 with both processes running as they usually do (Search with higher priority)
2) Core OS - Both processes running each in a separate rkt container (search with higher priority).
I saw no change in CPU / Network / Disk access metrics and my throughput remained the same.
Please note though, in my case I do not have way too many microservices as the general usage is. Also I use host networking. I also had no need for orchestration services like Kubernetes / swarm etc.,
TLDR:; No change between running product in container vs no-container mode with host networking, minimal containers and no orchestration.
Fleet management tools easily create hundreds or thousands of new virtual machines and run them through a complex npm bootstrap, when all you wanted to do was edit a file in /etc and restart node.
https://www.bloomberg.com/news/2014-11-14/5-numbers-that-ill...
Are docker containers at least more power efficient than virtual machines ?
> "Results: In all cases, there was a statistically significant (t-test and Wilcoxon p<0.05) increase in energy consumption when running tests in Docker, mostly due to the performance of I/O system calls."
Since most people adopt the "ephemeral" notion for containers, there's probably a lot higher rate of wiping and recreating bits, especially during dev and test.
I assume though, that a new process was designed to create efficiency in some other area that is more significant than any increase in energy consumption.
http://www.theregister.co.uk/2017/05/05/docker_docks_wallets...
which I posted separately a little earlier.
However, the article is titled "How does Docker effect energy consumption."
It is about an experiment which aims to quantify the impact on performance, which is quite interesting.