Open-sourcing Clusterman, a cluster autoscaler for Kubernetes and Mesos
engineeringblog.yelp.com
engineeringblog.yelp.com
Yelp is currently migrating to Kubernetes which is why support for Kubernetes autoscaling was recently added to Clusterman.
Of course, we notoriously underestimate the costs of roll your own. So a fair and common question is: why did you roll your own? What cost was so compelling?
What sometimes happens though is that the roll-your-own decision, made at time A, is later critiqued based on the options available at time B. If the decision was being made at time B, then the roll-your-own decision might be a symptom of "Not Invented Here Syndrome". But accusing someone of NIHS without accounting for when the original decision was made is unfair.
Hence my joke name, "Not Invented Yet Syndrome". Why didn't you use the alternative? Because it didn't exist or wasn't applicable.
Just wanted to clarify that running Cassandra & Kafka on Mesos is much easier now with the DC/OS Commons SDK.[1] Spark has always been supported on Mesos.[2]
[1]https://github.com/mesosphere/dcos-commons
Cassandra:https://docs.d2iq.com/mesosphere/dcos/services/cassandra/2.7...
Kafka:https://docs.d2iq.com/mesosphere/dcos/services/kafka/2.8.0-2...
[2]https://docs.d2iq.com/mesosphere/dcos/services/spark/2.9.0-2...
In fact DC/OS includes open source Mesos with no modifications, what they add is packaged bunch of different API providers running as Mesos apps: Marathon, DNS, etc., and also a somewhat dumb installer.
Auto-scaling was long possible for Marathon/Mesos, and is described at DC/OS website, though does not involve any of DC/OS additionals: https://docs.d2iq.com/mesosphere/dcos/2.0/tutorials/autoscal... .
In my org, we run open source Marathon/Mesos deployments for years without any DC/OS additionals, and also a slightly modified versions of tooling below:
https://github.com/mesosphere/marathon-autoscale https://github.com/mesosphere/marathon-lb-autoscale
There are also available several open-source auto-scalers for Marathon/Mesos apart from these two.
This new one from Yelp seems to be promising, since it's battle tested within organization which has many kinds of different workloads, so I would give it a try for some new cluster.
Side question: I know in the article it was mentioned that they liked their signal approach to scaling because it allows them to preemptively scale. I'm just not sure why that wouldn't be achievable by scaling the replica counts of your deployments based on signals.
If scaling happens based on pending pods then just scale your pods so that they're configured to handle your predicted traffic. Then the cluster will obtain your desired state. Am I missing something?
If you scale just the number of pods, if there isn't enough available capacity in the cluster, they are unable to deploy until that capacity brought online. By emitting a signal to increase the number of nodes in the cluster just before we think that capacity is about to be needed, we can ensure that the new pods are launched near instantly when the deployment is actually scaled up. This is mostly useful when you're scaling by large increments (i.e. hundreds of pods) that far exceed the spare capacity available in your cluster.
or maybe catching wind of some dev keys that really are root keys..
many reasons to sanitize git history before open sourcing. in fact many organizations i have worked with still maintain two separate repos, one internal and one open source using fancy magic (either with git or with additional tools) to sanitize and sync commits between the two. i've seen code commits to a large organization that are then packaged up and inspected for license and security violations in an untrusted environment.. many reasons to keep two (or more) running copies
git commit -am "hope this works"git commit -am "i did a stupid"
https://news.ycombinator.com/newsguidelines.html
Btw, that rule doesn't mean tangential topics aren't important—often they're more important than the thread topic. We have the rule to prevent interesting/unusual discussions from getting supplanted by boring/predictable ones.
I think engineers should be considered dis-associated.
But nobody said they are somehow immune from doing bad things. So nice strawman.
Why are german soldiers held in a negative light even though the war decisions were by its commanders?