730 karma · joined July 26, 2011
http://mesosphere.com/2015/02/26/deploying-with-docker-swarm...
In the Mesosphere world, Kubernetes is a "datacenter service" which is installed on your datacenter so that you can run Kubernetes workloads. You might also want to install DEIS to run DEIS-organized workloads. Or Spark, for Spark workloads... and so on -- all multitenant in the cluster. This is what the DCOS is uniquely good at, and why it qualifies as a true operating system.
The Mesosphere DCOS is built around the Apache Mesos kernel
The Mesos kernel was developed at UC Berkeley in 2009 [1].
Spark was written as a sample app on top of it [2].
Ben Hindman and his colleagues at the UC Berkeley AmpLab had always envisioned Mesos as a kernel inside of a full-blown operating system [3]. They finally brought it to market.
[1] https://www.usenix.org/legacy/event/nsdi11/tech/full_papers/...
[2] "We have implemented Mesos in 10,000 lines of C++. The system scales to 50,000 (emulated) nodes and uses ZooKeeper for fault tolerance. To evaluate Mesos, we have ported three cluster computing systems to run over it: Hadoop, MPI, and the Torque batch scheduler. To validate our hypothesis that specialized frameworks provide value over general ones, we have also built a new framework on top of Mesos called Spark, optimized for iterative jobs where a dataset is reused in many parallel operations, and shown that Spark can outperform Hadoop by 10x in iterative machine learning workloads." ibid.
[3] http://people.csail.mit.edu/matei/papers/2011/hotcloud_datac...
[1] http://www.cs.berkeley.edu/~rxin/db-papers/WarehouseScaleCom...
When you remove the VM smokescreen and count physical boxes it's more like 1 person/100 machines, which is abysmal. I've seen order-of-magnitude people efficiency increases with automation like we're discussing here.
My next startup will be built on Mesosphere. Faster time to MVP and no "go dark for 18 months" when I have to scale.
Here is a short tutorial for standing up Mesos on a single CoreOS instance: https://mesosphere.com/docs/tutorials/mesosphere-on-a-single...
Also, locality and latency can both be expressed in terms of placement rules and schedulers on Mesos can use those rules to guarantee or express preference for task placement that optimizes around reduced latencies.
The point that I take away is that the ability to express your needs in a declarative way (e.g., "place these two tasks such that they have such-and-such latency) is much more scalable, flexible and resilient than coding to machine-specific internals. The latter is easier to update and supports delegation of responsibilities.
John Wilkes of Google puts it this way:
"Our own experience has been that allowing our developers unfettered access to the internals of infrastructure systems has been a problem, and we're moving away from that model as fast as we can.
Constructing large-scale complex systems with many interdependencies leads to brittle, fragile systems if they rely on internal implementation mechanisms.
Allowing internal customers to rely on internal implementation mechanisms has made it hard to adopt new technologies, because we only know what knobs they set - not why.
The fix for both is similar: describe the desired end state, not how to get there."