Weave – The Docker Network
github.com
github.com
Take a look at GRE and/or VXLAN and the kernels multiple routing table support. (This is precisely why network namespaces are so badass btw). Feel free to ping me if you are working on one of these and want some pointers on how to go about integrating more deeply with the kernel.
It's worth mentioning these protocols also have reasonable hardware offload support, unlike custom protocols implemented on UDP/TCP.
Edit: Hi Joseph! Just realized it was you :)
nolan@cumulusnetworks.com
Unless someone did something crazy, traversing a veth pair should just be doing a little bookkeeping on the SKB, no data copies at all.
However I have strong doubts about the network performance, not only the overhead of the UDP encapsulation (that should be quite small), but mostly the capturing of packets with pcap and then handling them in user-mode. Looks like a lot of context-switches, copying and parsing with non-optimal code paths. Are there any benchmarks available?
My feeling is that this will consume large amounts of CPU for moderate network loads and thus be unusable with most NoSQL kind of systems that benefit from clustering across hosts?
We've got some issues filed to look at pcap alternatives and also generally aim to improve performance.
re suitability for NoSQL clustering... depends on where the bottlenecks are; if you want to cluster for HA rather than scale, i.e. there aren't any real bottlenecks, then weave will work well. Same if you want to cluster because of CPU or memory bottlenecks. If, otoh, networking is the bottleneck then adding weave into the mix isn't going to improve matters.
To retain the essence of how weave operates, this would likely not just be complex but impossible, short of kernel hackery.
> you would really only need to do processing at the beginning of a connection and for some ARP request
Weave needs to look at every Ethernet packet. Well, the headers at least. It's a virtual Ethernet switch. It doesn't even really know about IP, let alone TCP streams. See https://github.com/zettio/weave#how-does-it-work
Right now if you run a Docker container and run something within it as root, if that application gets compromised someone can break out of that container and alter the master (due to the way Docker links container-root and system-root, a root user in a container is effectively a root user on the whole system).
Docker are working on allowing containers to run entirely in user-mode (thanks to improvements in LXC). This would mean that you can run a process as root within a container, and if that gets compromised there is near zero chance of leveraging that into damaging the master OS (since it will just have normal user privileges).
Here's an article about their progress (to usermode):
http://s3hh.wordpress.com/2013/07/19/creating-and-using-cont...
To quote Docker's own documentation[1]:
> However, it has been pointed out that if a kernel vulnerability allows arbitrary code execution, it will probably allow to break out of a container — but not out of a virtual machine.
In other words, right now, a root process is likely able to escape a Docker container. You can use SELinux, AppArmor, and similar to somewhat mitigate that when it happens but neither are near as powerful as having that usermode isolation on the master.
If Docker is able to get usermode containers working, it will be very difficult for a Docker container to either alter other Docker containers or the master system (other than over the network, maybe).
[1]http://blog.docker.com/2013/08/containers-docker-how-secure-...
Because docker removes all the trouble of running applications that you need for your development: databases, application servers, queues...
I love the fact that I can focus on my code and not on all those details that stole so much of my time.
Sure, you can run MySQL in Docker, but it's a far cry from running it on native xfs with aligned partitions and whatever fancy you feel configuring. And since docker containers are very reusable, whereas backing data is by default should be persistent, my impression is that it's too easy to accidentally remove a docker container.
Generally, you want to be able to answer these questions when it comes to operating your databases:
What are the failure points? What is the impact of each failure point? What are the SINGLE points of failure? What is my recovery pattern? What is my upgrade experience? What is the operational overhead in the applications running ON the product? What is my DR strategy? What is my HA strategy?
Pure Docker, and no other tool that we are aware of in the docker/container ecosystem that we are aware of provide really good answers to these questions when it comes to databases. That is what I mean when I said running databases in containers in a nightmare. It is possible today, for sure, but it is extremely complex operationally and its why it is so rare to see prod databases running in production today.
This will be persistent, and will survive when you destroy the container. I use this to e.g. share a /home directory between a dozen experimental dev container I use to run my various projects - each container ensures I keep track of exact dependencies for each individual project, while I get to have a nice "comfortable" swiss-army-knife container with my dev tools and all all project files.
I also run a number of database containers which use volumes where I bind mount host directories to ensure persistence so I can wipe and rebuild the containers themselves without worrying about touching data.
I run a few MongoDBs with volumes, but I'm not confident that I won't accidentally start two with the same volume, or that someone won't accidentally delete the volume, or .. or .. or.
As I've written to a sibling comment, I don't consider it a hard problem, but it hasn't been taken care of .. yet!
And do you simply define the container as a volume to ensure it stays persistent? That was the feeling I got from the docs, but again, might just be flagging how little I know at the minute...
docker run --volumes-from=my-data app-beta-container
docker run --volumes-from=my-data app-prod-container
That would share the data store, however the real way you'd do this would be....
docker run --name=my-data -v /host/data:/container/data data-image
docker run --volumes-from my-data --name my-database database-image
docker run --link=my-database beta-app
docker run --link=my-database prod-app
Doing --link will allow those two containers to network-communicate and you should only be communicating with your database over the network anyways.
You don't need to make a container persistent: Docker, by design, will never remove anything unless you explicitly ask it to. If you want to separate the lifecycle of a directory within your container, so that it stays behind after you explicitly remove the container, or to share it between containers - that's when volumes are useful.
It looks like volumes have evolved significantly since the feature was introduced, you might want these links, sorry I haven't reviewed them myself:
https://docs.docker.com/userguide/dockervolumes/
http://crosbymichael.com/advanced-docker-volumes.html
(I actually do keep my backing data in the containers, we have institutionalized backups where all of the important data is already kept in git anyway, so instance clones are in fact disposable for me even though they have all of the important backing data in them.)
On a docker host you have the docker daemon, and whatever auxiliary stuff you need to orchestrate either the containers or the host (update docker itself, and so on), you have space for /var/lib/docker, and that's it. Volumes are always somewhere on /host/data. That means you have to make up a scheme and convention, and cook up scripts and add it to your already quite dynamic mental model.
If you go and want to manage volumes, you need something for that. And currently everyone and their cats have their own solutions (because there is one they claim to use and one they use, and one they hack on to use later). I'm not claiming it's a hard problem, just that it's not taken care of yet.
Maybe Flocker will deliver, I haven't checked it since it was posted 5 minutes ago :)
I'm not saying that vagrant > docker. The way I see it, docker is great if your infrastructure is using it all the way. If your prod setup if not dockerized, using docker in dev seems to me counterproductive than spinning up a VM and provisionning it with ansible or puppet to achieve production replication. As @netcraft said, I don't see why I should "change my server architecture" to use docker in dev.
I totally agree that startup time of a container is far less than a VM, but I don't see how docker "removes all the trouble of running applications that you need for your development: databases, application servers, queues"
You still need to install, configure these services, make sure that the containers can talk to each other in a reliable and secure way, etc.
That said, all of those fiddly library dependencies are where i struggle the most at work. If i could just build a docker image and hand that off, it would save me a lot of grief with regard to getting deployment machines just right.
I do have a great deal of experience with legacy environments, and it seems like the only way to actually solve problems is to run as much as possible on my machine. Lowering that overhead would be valuable. Debugging simple database interaction is fine on a shared dev machine. a weblogic server that updates oracle that's polled by some random server that kicks of a shell script... ugh. Even worse when you can't log into those machines and inspect what a dev did years ago.
If you've got a clean environment, there's probably not as much value to you.
In the end, it's just so slow that nobody uses it locally. Even on a beefy Macbook Pro, spinning up the six VMs it needs takes nearly 20 minutes.
We're looking at moving towards docker, both for local use and production, and so far I'm excited by what I've seen but multi-host use still needs work. I'm evaluating CoreOS at the moment and I'm hopeful about it.
* Install your stack from scratch in 6 VMs: slow * Install your stack from scratch via 6 Dockerfilea: slow * Download prebuild vagrant boxes with your stack installed: faster * Download prebuilt docker images with your stack installed: fastest
The main drawback of Vagrant is that afaik it has to download the entire box each time instead of fetching just the delta. That may not matter much on a fast network.
On the last project where I had to regularly run many VMs on my laptop, the software being tested used more than 1GB. Calling it 1GB total per VM and sticking with the 53MB overhead, switching to containers would have reduced memory usage by 5%. Again, to my mind that's trivial.
We've created weave networks spanning hosts on EC2, GCE and local data centres.
I suppose service discovery is out-of-scope for this project but having some sort of weave-wide hostsfile would certainly simplify it. Am I misunderstanding the project?
Meanwhile, two points of note:
1) In weave the IP addresses can be much "stickier" than in other network setups, i.e. a moving a container from one host to another can retain the containers IP. That means it is quite amenable to relatively static name resolution configurations, e.g. via /etc/hosts files.
2) Since weave creates a fully-fledged L2 Ethernet network between app containers, name resolution technologies like mDNS that rely on multicast should work just fine.
So, in summary, while weave currently does not have any built-in service discovery, existing solutions and technologies for that should be relatively easy to deploy inside weave application networks, until weave itself grows these capabilities.
Also Rudder is Layer 3 and Weave is Layer 2 and Weave can encrypt traffic.