Rudder: An etcd backed overlay network for containers
coreos.com
coreos.com
Initially we wanted to use an existing in-Kernel encapsulation format like a simple ip-ip encapsulation. However, IP-IP doesn't work on AWS. Then we looked at VXLAN but it relies on multicast which doesn't work on most cloud networks either. Most recently we started looking at the VXLAN DOVE extensions and are getting a prototype together for this.
tl;dr the initial goal is to show that something generic is needed and can work, we will get something that is performant and/or has encryption next.
What I'm really excited for are the possibilities of docker containers with public-routable IPv6 addresses. It would move the world away from "one host: many services on different arbitrary ports", and back to the "one host: one service, possibly speaking a few protocols with ports being used for OSI-layer-5/6 protocol discovery" model of the 1970s (and eliminate the madness of SRV records, besides.)
Imagine if, say, bitcoind (which normally speaks "JSON-RPC" to clients -- a specific layer-6 encoding over HTTP) sat on "bitcoind.host:80" instead of "host:8332". Suddenly, it'd be immediately clear to protocol clients (e.g. web browsers) which hosts they could or couldn't speak to, based on the port alone! The whole redundancy between schema and port in URLs could go away: they'd be synonymous. And so on.
If there's NAT between the hosts (typically this is the situation at home) then Rudder will not work. We may add limited support for NAT, e.g. ability to specify a public IP to use instead of the one assigned to the NIC. However to make it easy to run from home would require doing NAT punching like STUN. We're focusing on making this useful for running in the datacenters with Kubernetes clusters.
Things are not as easy on other cloud providers where a host cannot get an entire subnet to itself. Rudder aims to solve this problem by creating an overlay mesh network that provisions a subnet to each server. ... is unclear.
What host for virtualized infrastructure needs an entire, fake, non-internet-routable subnet that it cannot provision itself?
I believe there's a broken one size fits all network architectural assumption or provisioning methodology at the root of all this.
(Edit as reply to child as rate-limited: Sounds like I was right, and it's docker's fault. How is this not better solved with the standard approach of applying network namespaces and/or unique interfaces to containers?)
Kubernetes has an idea of a pod -- a group of containers that share an netns and have an IP.
Reasons you might want a pod: * A thick client or client side proxy that follows the ambassador pattern for service discovery and access. * A data-loader and data-server pair. The loader would grab data from some persistent source and write it to disk or a shared memory segment. The data-server would then use that data and serve it up. You'd could run the data-loader at a lower QoS so it doesn't stall the data-server. * Some sort of server and a log saver. The log saver could periodically batch up and compress structured log data and upload it to a persistent store (such as BigQuery in GCP). You want to build/configure/restart/upgrade the log saver separately from the server. You'd also run the log saver at a lower QoS.
Inside of Google we have all sorts of examples where we have sets of containers/tasks/processes that are co-scheduled onto a machine and work together.