Swarm vs. Fleet vs. Kubernetes vs. Mesos
radar.oreilly.com
radar.oreilly.com
* persistent volumes: let frameworks have directories on agent machines returned to them after an availability event happens, which is nice for replicated stateful services
* maintenance primitives: schedule machines to go offline at certain times, and ask frameworks if it’s safe to take nodes offline. This will soon start being used for stateful services so that they can vote on when it’s safe to take out a replica, and to trigger proactive rereplication when maintenance is desired.
* oversubscription: if you have an agent that has given away all of its resources, but the agent detects that there is still some unutilized CPU, it can start “revokable tasks” to fill up the slack up until it starts interfering with existing workloads.
For me, the cleverest part about Diego is that it turns placement into a distributed problem through an auctions mechanism. Other attempts at this focus on algorithmic approaches that assume perfect information. Diego instead asks cells to bid on incoming tasks and processes. The auction closes after a fixed time and the task or process is sent to the best bidder. This greatly reduces the need to have perfect central consistency in order to perform bin-packing optimisation -- in a real environment, that turns out to matter a lot.
Cloud Foundry is large and very featuresome, so for those who want a more approachable way to play with Diego, try Lattice[2].
Disclaimer: I work in Pivotal Labs, a division of Pivotal. Pivotal is the major donor of engineering effort to Cloud Foundry. I worked on the Buildpacks team and I've deployed commercial apps to PWS, so I'm obviously a one-eyed fan.
[1] https://github.com/cloudfoundry-incubator/diego-design-notes
Diego works by auctioning tasks/LRPs to cells, which bid based on free resources.
Mesos works by auctioning free resources to tasks/LRPs, which accept or reject based on their own policies.
The analogy I got from Onsi was that Diego is "demand-driven", Mesos is "supply-driven".
They have different sweet spots. Diego is built to be the heart of a PaaS, so it aims for fast success or failure for accepting tasks and processes. Consumers can decide whether to build further scheduling intelligence, but Diego doesn't make it explicit.
Mesos exposes some of the placement/scheduling logic back to the client explicitly. For example, if you are OK with a gigantic batch job running intermittently, Mesos can more or less do that without fuss. The job will simply wait until enough free resources are available on the cluster. This is one of the observations that drove Mesos. Whereas Diego will tell you very quickly if that capability is unavailable currently.
On the flipside, Diego requires no additional configuration. It also doesn't require a central view of available resources to be maintained in order to perform placement. That said, Diego does maintain a central view of tasks and LRPs that have been placed, in order to relaunch failed ones.
That said, I have sometimes been wrong about my understanding of Diego and Mesos, so YMMV.
I should also note in passing that Diego can place and manage Docker containers, Windows .NET apps and classic buildpack-staged apps, with the architectural flexibility to extend to more kinds of managed resources (eg, with the right backend, it could manage unikernel apps). That alone sets it apart from the pack.
State is handled by services bound to your application. CF injects the connection details into the staging and runtime environments. Applications using buildpacks generally have that connection autoconfigured.
For example, Rails apps have database credentials autoconfigured if you have a database bound. Spring apps get a DataSource. And so on.
It looks like k8s has a way to handle volumes, so maybe I'll have to play with that.
But you'd be surprised how many apps can be ported to some semblance of cloud nativeness. Most of the stuff that Absolutely Must Be Stateful can usually adapted to something else. Sticky sessions go to Redis or Memcache, local files can often be relocated into a shared filesystem/blobstore like Riak CS etc.
The big thing about PaaSes is not that they make it a smooth, double-percent-ROI incremental improvement from the status quo. What happens is that unlocking true continuous deployment capability unblocks process backpressure throughout the entire organisation. The value you get from a complete shift in how you develop software swamps the cost of carrying legacy until it can be replaced.
Putting it another way: the sucky parts are tactical. The awesome parts are strategic.
[1] http://blog.pivotal.io/pivotal-cloud-foundry/case-studies/th...
For fleet a timer is just another systemd unit - there's no special support for them -, so you get a simple "cluster wide cron" pretty much for free.
Fleet works fairly well, though it does have some minor maturity issues (I've gotten it into weird states when machines have left the cluster abruptly and rejoined where e.g. it refuses to schedule some unit afterwards; solution: change the name of the unit) . No idea ho well it'll scale in larger deployments.
You are actually better of managing each server individually than using swarm.
(hard earned truth)
I'm still wondering if and how they're going to address the idea of co-scheduling groups of containers, like pods in Kubernetes. This is a common need, but Docker don't seem very keen: https://github.com/docker/docker/issues/8781
Disclaimer: Co-founder of Rancher Labs and chief geeky guy behind Rancher
Also, I want to make a plugin/service that's dependent on what you've build solely for the pun "Huevos RancherOS"
At least, last time I checked. I can see components in the Marathon API to help you build such a thing. But Kubernetes deals with a fundamentally more abstract set of primitives: groups of containers (pods), replicated pods, and automatically managed dispatch points to said groups called services.
The Marathon documentation literally recommends you hand-roll a periodically updated haproxy service to accomplish that last bit. While K8S introduces a fair amount of overhead to accomplish this feat reliably, that overhead pays off by having a very responsive and available system.
So yeah. K8s is working one level up the stack, as I see it. I think you can make something reasonable with only Marathon, and that thought is backed up by the reality of people doing it. But my own experience suggests K8S does more for you here. Hence my excitement about K8S-mesos.
It listens for marathon callbacks (for app create/update/scale), or when a task is killed/lost and updates the backends in etcd, propagating to all vulcand load balancers.
The tool is available on github [2]
[0] https://vulcand.io/proxy.html#backends-and-servers
[1] https://mesosphere.github.io/marathon/docs/event-bus.html
For a small number of nodes (< 100) and requests (< 8000 req/s on each vulcand server) this vulcand approach is ok. For large scale, HAProxy/Nginx supports a lot more req/s than vulcand, and I think it can also be configured using confd [0] using a similar aproach: listen for marathon events, update etcd keys, then confd will listen for changes and reload HAProxy/Nginx
[0] http://ox86.tumblr.com/post/90554410668/easy-scaling-with-do...
https://github.com/mesosphere/marathon/blob/master/bin/hapro...
I'm not sure what you mean by "hosting external services reliably" - what's external and who is unreliable?
There's nothing wrong with these decisions. It's just odd to compare this group. It makes me wonder how people are considering these options and in what context.
You cannot field a reliable and available service to the outside world without this component. And the instrumentation and operationalization of this component is tricky.
I don't get why you think I can't compare that group though. What should I compare Kubernetes with? The group represents the various approaches to clustering and orchestration that most people will consider for container deployments. They are all slightly different, which is what makes for an interesting comparison.
In addition, of course, the task of learning, implementing, and evaluating the options take a large amount of time on top of the time we already spend (mostly) manually maintaining infrastructure.
Articles like this are a great stepping stone.
(disclosure, I'm a Juju dev)
If you are running OSX, Windows, or a non-Ubuntu linux, you'd need an Ubuntu VM to run the local provider (we support Centos, Windows, and OSX for the juju client in general to control remote providers like Azure, Amazon, etc, just not for running the local provider).
We've found the Kubernetes primitives to be the easiest and most straight-forward to work with while still providing a very powerful API around which to wrap all sorts of custom tooling.
There's also technologies like flocker that will allow you do this with docker volumes (which can then be run inside mesos).
Both of these are in a pretty early stage of life, so there's some rough edges, but if you really need it, it's there.
Ceph itself seems really stable, though we did have an issue with kubernetes not aquiring a lock on ceph and thus having 2 nodes write in the event that one went unresponsive. This was patched in a later version, but make sure it's merged into the release you go with.
At Quobyte (http://www.quobyte.com, disclaimer: I am one of the founders), we have a built a fully fault-tolerant distributed file system. This allows concurrent scalable shared access to file systems from any number of hosts. Think of a /data that is accessible from any host and can be mapped in any container.
What I found pretty neat was that we could easily do a mysql HA setup on Mesos: put mysql in a container, use a directory on /quobyte for its data, and enable Quobyte's mandatory file locking. When you kill the container, or switch off/unplug its host, the container gets rescheduled and recovers from the shared file system.
Kubernetes offers this as a plug-in today.
https://github.com/kubernetes/kubernetes/tree/master/example...
Would people agree?
If you have a very large cluster (1000s of machines), Mesos may well be the best fit, as you are likely to want Mesos's support for diverse workloads and the extra assurance of the comparative maturity of the project.
The big difference with Kubernetes is that it enforces a certain application style; you need to understand the various concepts (pods, labels, replication controllers) and build your applications to work with those in mind. I haven't seen any figures, but I would expect Kubernetes to scale well for the majority of projects.
I'm not sure what the trade-offs are with running Kubernetes on Mesos - that would make for an interesting article.
Running workload-specific schedulers like Kubernetes on Mesos is one of its fundamental ideas. From the paper:
"It seems clear that new cluster computing frameworks will continue to emerge, and that no framework will be optimal for all applications. Therefore, organizations will want to run multiple frameworks in the same cluster, picking the best one for each application."
https://people.csail.mit.edu/matei/papers/2011/nsdi_mesos.pd...
You might also be interested in a talk I gave at the Kubernetes launch in August about our work coming the two: https://youtu.be/aXcdHwQ5GgQ
a1r is exactly right - Mesos does great in virtualizing your data center, Kubernetes is a framework on top of that.
Here's a blog discussing the current state of K8s scale: http://blog.kubernetes.io/2015/09/kubernetes-performance-mea...
Note that Bob Wise at Samsung has been driving some horizontal scale testing, and has got a 1000-node cluster up and running, so that's a current "best case" scale number.
(work at CoreOS)
In our experience, it isn't so much that Kubernetes enforces a certain style as that it doesn't force services to understand the scheduling layer. A Twelve-Factor-style app will deploy very nicely and easily into a Kubernetes cluster.
Or should, when the K8S-Mesos stuff is a bit more mature.
What I really need is a "distributed cron", first and foremost, with the additional requirements of: it being lightweight on resources, and it being multiplatform. Dkron seems to fit the description pretty well, I'm looking for any feedback from users.
They are looking for services like e-mail (SMTP, IMAP, maybe webmail), website (static + maybe wordpress), source code hosting (Subversion) and reviews (?), CI (Jenkins?), devops, and CRM...
For the environment conjured in my head by what you're describing, I would probably use something like Salt or Puppet to drive installation, config management/monitoring and upgrades, possibly on top of oVirt or similar.
I don't see an advantage for tiny shops containerizing everything at this point, at least until some bright shiny future where containerization is much more of the norm and platforms like this are much more mature and easy to manage. Your sysadmin almost certainly has better things to do than adding wrappers and indirection to a tiny environment.
I haven't installed Kubernetes directly, but I have setup and used Openshift v3 which adds a PaaS solution on top of Kubernetes. Setting it up is really easy, and they've release an all-in-one VM to demo it out: http://www.openshift.org/vm/
The other option is to use a hosted solution. Google Container Engine (https://cloud.google.com/container-engine/) is essentially hosted Kubernetes.
PS I work for Red Hat. OpenShift V3 is our product, and we contribute a lot to Kubernetes.
Rolling your own is very time-consuming if your goal is to deliver simple apps.
(disclosure: I am a cofounder of Flynn)
Why is this the case? My understanding of containers is that they solve exactly this problem - restart, rescale only a fraction of the application (only an individual service).
If all of a cluster is deployed to a single host, and that host dies, then the recovery time is the time needed to reschedule the container elsewhere, load it onto the machine (if not cached), and execute it. This is all downtime, since every instance was lost.
If instances are on more than one host, then even in the case of a single host's failure the application is still available while the failed instance is rescheduled and started.
Many don't, of course...
If you have a reachable etcd cluster (which you can run on a single machine, though you shouldn't) and your machines run systemd, and you have key'd ssh access between them, you can pretty much start typing "fleetctl start/stop/status" instead of "systemctl start/stop/status" and feed fleet systemd units and have your jobs spread over your servers with pretty much no effort.
For me it's an effortless way to get failover.
E.g. "fleetctl start someappserver@{1..3}" to start three instances of someappserver@.service. With the right constraint in the service file they'll be guaranteed to end up on different machines.
Having a set of uniform nodes where you schedule containers is nicer, in my opinion, than managing an infrastructure where you've scripted applications to go places.
Sure, you could make puppet manage containers on uniform nodes but then we're having a massive convention vs configuration argument, which we shouldn't bother with because the distributed schedules are giving us a lot of important stuff for mostly free.