Introduction to Apache Mesos
antonlindstrom.com
antonlindstrom.com
We have a separate, much smaller cluster to test new Mesos (and Chronos and Marathon) versions. But the distinctions "production", "staging" and "dev" have become much more nuanced, so we've settled on discussing the "application" environment versus the "infrastructure" environment. Much as, as a startup on AWS, you wouldn't distinguish the RDS instance of your database (e.g. 9.3.1), you would distinguish the version of your database on an RDS instance, we distinguish the versions of our apps on the production cluster and not the version of the production cluster. One of the team was an ex-Googler and he said that Google did much the same.
The one thing Marathon and Chronos currently lack is a prioritization mechanism so we're building that as a Chronos task that monitors and scales down/up Marathon tasks by their priority (as represented in their id or tag).
Is there some common phrase relating belching and being happy that I haven't heard?
I don't think that it's backwards compared to other schedulers (many of them are monolithic) but I think it can be improved by using transactions and optimistic concurrency control. This is discussed in the Omega paper and I think there's work to integrate the shared state scheduler into Mesos.
Mesos decides how many resources to offer each framework, while frameworks decide which resources to accept and which computations to run on them.
I still can't find documentation about what fairness policies are actually implemented, though.
With the current design, if I want to launch a Docker container, I have to use its scheduler, even though its scheduler might not be as good as some other resource's scheduler. And I might have to accept a regression in the scheduler in order to benefit from a bugfix in the executor.
I'm running Elasticsearch nodes on a dedicated Mesos slaves and I’m still not sure how much memory should I allocate to Elasticsearch task e.g. all available memory as reported by Mesos or some smaller amount to leave something for the system? Please note that I'm asking about memory allocated to the Marathon task not about JVM heap size.
AKAIK, there isn't a right way to run databases because the persistent storage layer into mesos is still being baked.
That's pretty important point, would have thought persistent storage would have been first point to get working for a project like Mesos. Also, their homepage ( http://mesos.apache.org/ ) outright states: "Apache Mesos abstracts CPU, memory, storage, and other compute resources away from machines [...]". What storage are they talking about if not persistent storage?
Essentially the current mesos disk quotas are just so that tasks that run, run with a minimum amount of disk space. I think what they are trying to accomplish in future releases is some way you could build an EBS Style disk space management on top of mesos.
https://issues.apache.org/jira/browse/MESOS-2018
With a first draft of the user documentation at:
https://gist.github.com/mpark/e8ee4eb9671bdb252c4f
It will be really slick once this makes it all into Mesos 0.23
With that said, I often wonder how many people are using Mesos/Marathon before they have any need for it? Using it on for 4 hosts vs. 40 or 400?
Anything more than a single host requires additional co-ordination - picking Mesos is mostly no different than rolling alternate methods for doing so.
It's not choice with the least overhead, but it's a possibility, for those for whom it has the right conveniences...
As someone working on scaling microservices, I keep being disappointed by potentially useful services that turn out to require a JVM-based language such as Java or Scala. For example, Kafka looks very decent, but the high-level client is written in Java; if you're not on the JVM, you're stuck implementing a lot of the client yourself. As far as I know (from the last time I looked at this stuff), the Zookeeper client is similar, whereas Spark and Storm both require that you write processing code on the JVM, and libhdfs is apparently still JNI-based, not native.
For someone using Docker, is there anything competing with Mesos that isn't wedded to Javaland?
Not at all. My company has a cluster with a few hundred Slaves and we're mostly a Python shop, with some C++ for machine vision.
Mesos certainly has its issues (as does everything), but its awfully nice for micro-services: if you can package your service into a Docker container, then you can launch it into the cluster and Chronos/Marathon/Mesos will take care of making sure that it's run/running.
I missed the deadline for the Mesos conference (I was unaware of the deadline), but I'm trying to squeak in a talk about "Using Mesos at [small] Scale" because we're a small company and Mesos has allowed us to do a bunch of big company stuff.
>is there anything competing with Mesos that isn't wedded to Javaland?
Yes, I can look at the source code and see that it uses te JVM (Chronos uses Scala), but, AFAICT, Mesos isn't "wedded" to anything. All of the components are API-driven. I apt-get install it, I run it, I send jobs to it, it works and it behaves well. Better I can poke at the APIs of any of the services to find out what is happening. So we use Marathon for service-discovery and run Chronos, a framework, under Marathon. Makes finding Chronos, which could be on one of 200 machines, quick and easy.
I have to say I'm wary of big frameworks like this that insert themselves as a kind of monolithic control structure for everything.
My ideal setup is always one where I pick and mix the best modules for the job, and where I can write some interface glue to let my apps slot into the system, as opposed to writing my apps for a specific API (as tends to be the case with, say, Hadoop, and which of course would tie the whole platform to that API, making it hard to migrate to something else).
Sounds like Mesos is pretty modular and open in that respect?
Lattice[0].
It's built from a few components extracted out of Cloud Foundry, all written in Go. In particular it includes Diego, an orchestrator. It can be run locally, on AWS, Google and DigitalOcean.
In terms of the key architectural difference, Onsi Fakhouri explained it to me this way: Mesos is supply-driven, Diego is demand-driven. Mesos keeps a list of pending workloads and waits for a report of available resources that fits. Diego instead receives a request for resources and then stages an auction amongst available cells.
Disclaimer: I have worked on Cloud Foundry (on the Buildpacks team here in NYC) and I am employed by Pivotal, we do the lion's share of the work on Cloud Foundry.
Between Exhibitor and Curator, I honestly find Zookeeper so straightforward and easy to work with that I don't quite understand the popularity of etcd.
I'm not saying etcd won't ever overshadow Zookeeper, it probably will with the momentum behind it, but as an ops guys, I wasn't willing to bet production application service discovery on it.
Also there was that one time they changed how cloud-config was parsed, so if "#cloud-config" wasn't on the very first line without preceeding spaces, initialisation would fail. That was when I switched the reboot strategy to manual.
Tools like etcd and consul fit into the Unix philosophy of small, composable tools. Zookeeper is more a part of the Enterprise Java philosophy which many people have written off for various reasons, both rational and irrational.
Having run Zookeeper in the past and now having run Consul in production for the past ~6 months, I can't imagine ever running Zookeeper again, unless I'm using a tool that's built on top of it. Consul is just easier to use/maintain and we've yet to run into any problem with it. Zero problems in six months. In all the time I ran Zookeeper, I could never say that.
I mean, Consul is fine for what it does. I've used both, whatever. But if you're going to have something that does a bunch of things, I'd much rather have the one that supports the primitives to do what somebody needs, rather than trying to do it all itself.
(Personally, after trying to work with Terraform, I don't much trust Hashicorp's attempts to write code I have to rely on to work correctly and never ever break. YMMV, of course.)
- Single binary executable
- Compatible with ps (Zookeeper has the traditional java problem of showing up as .../bin/java followed by 4 lines of classpath)
- Arguments to consul don't need to be prefaced with -D (another common java problem)
- Passing -h to consul actually helps you figure out how to run it.
Oh, and the download is 1/3 the size of Zookeeper, and the executable includes the Go runtime whereas Zookeeper's java runtime is separate.
Etcd is actually much closer to the Unix philosophy. Consul seems to go more in the direction of similar Go tools like Docker where it bundles related activities together into one executable. But, then again, parts of the Unix ecosystem do this to (openssl, for one).