Docker Containers at Scale: Our Take on Docker Swarm
mesosphere.com
mesosphere.com
Even Exhibitor is not a silver bullet for managing ZK. It helps with rolling restarts and lets you do a sort of MySQL-esque bin-log replay of your state, but even with Exhibitor, it's not going to make for a smooth recovery at 4:00AM when PagerDuty rings.
The underlying design of how Mesos and Marathon communicate create situations where they have differing views of the state of the cluster, and you end up with "orphans" which are tasks that Mesos is aware of but Marathon knows nothing about.
In my opinion, the system feels like it was designed to be something else, and then all this functionality was tacked on later. Coincidentally that is exactly what the case is with Mesos. I think there is a lot of room for an improved user experience that these tools will struggle to provide.
There's a big naming and marketing confusion at the moment with the open-source project and the VC-funded startup having such a similar name. It's unfortunately giving a bad reputation to the Mesos project.
Curious why there actually isn't a trademark issue here...?
Twitter engineering recently blogged about Aurora with a bit more on its history: https://blog.twitter.com/2015/all-about-apache-aurora, and The New Stack recently published a two-part article on the project: http://thenewstack.io/twitters-aurora-relates-googles-borg-p..., http://thenewstack.io/twitters-aurora-replaces-operating-sys...
(full disclosure, I work at the tweet shop)
One big reason to use Marathon + Chronos + Mesos is to get multi-tenancy of compute and data workloads on your instances. You can't get that with web-service specific tech like Tutum.
[1] appuri.com
You can make an account and boot apache Spark in about 30 seconds using this link [0]. It's running in production right now for a lot of people, and you can run Mesos on top of Terminal if you want it [1].
Again, I'm not trying to push this on you if you're happy with how stuff works today, but I think we've made a PaaS that solves a lot of these problems (and we'll let you run it on your own metal too if you want it). Check it out at terminal.com if you have some free time.
[0]https://www.terminal.com/snapshot/c81e6215eba5799335a45b6936... [1]https://www.terminal.com/snapshot/44d4ee043422afec75dfd3bdaa...
It's kinda cool, or at least I think so.
What's your set up? How many zookeeper nodes do you run? What problems have you run into?
Many of the docs on etcd promote the use of the discovery service[1] as a convenient way of bootstrapping an etcd cluster -- really useful when you don't know the IP addresses of each cluster member up front. The discovery bootstrap method is also great for demos and testing environments, but as you correctly highlight this is not the ideal way to run a production setup.
With the release of etcd 2.0, I strongly recommend using the static bootstrap[2] method for provisioning an etcd cluster. The static bootstrap method provides the key documentation clues that help you reason about an etcd installation.
Finally, etcd 2.0 introduces support for bootstrapping using DNS[3], which provides the convenience of the discovery bootstrap, and the explicitness of the static bootstrap.
[1] https://coreos.com/docs/cluster-management/setup/cluster-dis...
[2] https://github.com/coreos/etcd/blob/master/Documentation/clu...
[3] https://github.com/coreos/etcd/blob/master/Documentation/clu...
We wound up proactively restarting the ZK cluster regularly, which improved stability.
Granted, it was our own software written to use it, and we suspect there were problems with the way it was written. It was easier for us to just rip it out than debug, however. I find it overly complicated to write against given the need for thick clients.
Consul phrases it well (https://consul.io/intro/vs/zookeeper.html):
"ZooKeeper provides ephemeral nodes which are K/V entries that are removed when a client disconnects. These are more sophisticated than a heartbeat system, but also have inherent scalability issues and add client side complexity. All clients must maintain active connections to the ZooKeeper servers, and perform keep-alives. Additionally, this requires "thick clients", which are difficult to write and often result in difficult to debug issues."
Edit: in retrospect our problems might have been solved by turning syncing off, as described here:
http://www.edwardcapriolo.com/roller/edwardcapriolo/entry/zo...
But you can figure out from the above how scalable Zookeeper is... not very. We run in physical datacenters and I certainly wouldn't be thrilled about building out snowflake RAID systems just for our ZK clusters (we generally try to use whitebox commodity hardware and we want individual nodes to be as disposable as possible).
It's ironic that ZK requires such consistency when the goals of Mesos are exactly the opposite.
Is it ironic? Seems to me that if you want to have a distributed cluster that can deal with worker failure and still be useful, you need to rely on something to durably maintain your state at a lower level.
No, individual nodes have gone down based on hardware problems. The system stays up. Jespen has given Zookeeper probably the most ringing endorsement of anything it's tested, so I don't know what you're on about.
AFAIK it allows you to replace ZK for etcd making it a lot easier to run this on top of coreos as a working etcd cluster is a fact in a working coreos cluster.