Data migration with Kubernetes and Flocker
clusterhq.com
clusterhq.com
I'd think that the workflow would be similar to migrating ZFS snapshots, just extracting volumes from docker directly, and sending tarballs to the new node, as opposed to the built-in ZFS send/receive. It probably wouldn't be as efficient, but it would make Flocker an easier sell to more people.
Another question - about what size do you think portable volumes need to be before they stop making sense and you need to move to some other kind of either shared storage (NFS) or remote object storage?
Doing a tarball based export-import backend would be an interesting Flocker backend ;)
With regards large volumes, adding more backends to Flocker will definitely open up the scope for volumes then a ZFS pool (zpool) you'd want to have on a single node. Generally though, when you need to scale data up to larger scales, it's better to use a distributed data store which can shard your data across a larger number of smaller stateful containers, which can also be managed by Flocker, but of course that's application-dependent.
I like how this migration works. What about live migrations? Maybe I misunderstood, but I think in the tutorial the service is offline for some time. What about maybe applying backpressure (tcp level) to timeout the request while migration is underway?
Edit: If we are doing tarball migrations, then maybe i'd rather have an rsync backend :)
Oh and if we are doing large amounts of data migration (gigabytes), maybe even use something akin to something like what bittorrent sync is doing? The scenario would be that you use bittorrent sync (this is a theoretical example..) to one-way migrate continuously (to maybe multiple hosts..??), then pause/hold traffic when satisfied, complete final sync and you are migrated.
I'm not sure this would even be possible outside of a virtual machine. VM hosts have significantly more information and control about the VM than the container host has about the container. For example, a VM host knows exactly what memory is hot and what is cold to enable sub one-second switchover times.
It might be possible to add such support to the Linux kernel itself to support process migration between hosts, but that's the level of work required.
The alternative that I was thinking about using your standard high-availability tools to create a new worker on the new host, then gradually remove existing workers on the old node. I think that might be the only way to really make "live" migrations work from a practical standpoint. For web-like services, this would work. For others, it may not be practical.
EDIT: Looks like the checkpointing running processes might work after all! (see: http://en.wikipedia.org/wiki/CRIU) I must admit that I'm still a little skeptical, but would love to see this working!
But, I wouldn't call that a "live migration" in the VM live migration sense.
However, I wouldn't worry too much about it, since you're absolutely right that most processes can restart gracefully.
This problem arises because Docker positions their intrinsics as novel containers, even though there was an entire field of prior art long before Docker showed up. Hence your confusion and worry checkpointing will never be supported; it already was, but you didn't know about it, because to you Docker == containers. One of my many problems with Docker, because they feed that.
Look how easy it is: http://criu.org/LXC
I have to agree though that it seems strange to me that Docker has gotten a lot of support so quickly when things like LXC, FreeBSD Jails, or Solaris Zones have been established for so long. I've only played with Jails a bit, having done more work with full virtual machines in Linux. The other containers, I at least had a familiarity with before Docker (nothing in production). However, I had not heard about CRIU until today. But it looks really interesting!
Regarding backpressure I am talking about reactive streams, of which I am a fan of late, this would be better suited perhaps at the application level [0].
Regarding infra/pod level migrations, I am wondering if this would be something to implement at the kubernetes level , perhaps a new type of controller/service ? You could then perhaps use labels to identify throttling/freezing/thawing of pods during migration.
The pace of Docker interoperability is frustratingly slow