For example, Openstack deployments are commonly automated using Puppet scripts. (Really amazing stuff if you haven't seen it before...)
DCOS may choose to reinvent the wheel, but unless there's a core innovation in the way they're handling orchestration or automation, I think they're better off if they leverage an existing battle-tested toolchain. Adding another tool will just mean more work for the (underworked and lazy?) DevOps teams who will be responsible for managing DCOS.
As someone who used to run a 6000 core farm in 2007 (its not 25,000) I can tell you that puppet isn't going to help task placement. It can create machine that will run a certain app, but without some heavy programming it'll never balance or detect need and respond sensibly.
These are all pretty similar: - configuration management: ensure package oracle-java-8 is installed on machines A,B,C with this specific configuration. - orchestration: ensure my-awesome-java-app on machines A,B,C is running to databases on machine D,E,F - deployment with constraints: ensure that four instances of my-awesome-java-app are running on at least 2 physical machines with over 4TB free disk space. - job runner: ensure that script X runs on a cluster every __ minutes. when script X runs, send the output to script Y
I think that you'll see task placement and job scheduling primitives being integrated into DevOps tooling in the next 6-12 months.
Saltstack already has many of the primitives in place for building out reactive infrastructure, http://docs.saltstack.com/en/latest/ref/runners/all/salt.run...
Would be great to hear about tools for Puppet, Chef, Ansible, other…
For distributed cron, we use jenkins. Which has the advantage of keeping "build history"
For task placement we use alfred (https://renderman.pixar.com/resources/RPS_13.5/rps_manuals/a...) yup its old. However it works like a champ, and its fast. (as in it'll dispatch thousands of tasks a second.)
You do hit a limit when you go over 6000 "slots" (each slot accepts one task, and the main dispatcher is single threaded). Dispatching is simple and task building has simple syntax that easily grows to thousands of tasks in one job. monitoring is also simple, as each task ships logs and exit status back to the dispatcher. It also has mechanisms to cope with bad/slow/unhappy machines.