Executing Cron Scripts Reliably at Scale
slack.engineering
slack.engineering
I find jumps like this hint at "political stiction", sometimes it's hard to get permission to do incremental updates to things, you have to wait until the smoke from the burning tires is unmissable, then get big political consensus and a "visible project" to allocate budget and time for what would otherwise be unsexy maintenance work.
However, I'm sure some folks would be tempted to add something like "designed and implemented a distributed task scheduler and execution engine for generalized asynchronous jobs utilized by X number of devs across Y teams" to their resumes.
Because I can't imagine why it would award that relevance. It's right there with "implemented function to reverse a list because the stdlib had a bug".
Literally went from a working "tidy little house" to building a massive sky-scraper.
Obviously Slack operates at a much larger scale than I do...but holy moly.
hey, good luck to them with that!
> The scheduling is approximate because there are certain circumstances where two Jobs might be created, or no Job might be created. Kubernetes tries to avoid those situations, but does not completely prevent them.
For example, if the task that runs your cron is down when your cron is supposed to run, then it won't run.
The slack blog says they did some tinkering like preventing nodes from going down at the top of a minute because that's when they think cron jobs are most likely to run. But at scale things are going to break when they break, and you have to weigh the pros and cons of designing the jobs to be robust to failure vs trying to organize failures to correspond to the needs of your jobs.
So I think there is space for solutions that make different tradeoffs. But it does seem vastly easier to tune an existing solution that someone else is maintaining than to build your own solution on top of Kafka.
[0] https://stackoverflow.com/questions/47691278/why-in-kubernet...
For a task runner, there are a lot of different behaviours you might want if the system crashes. Maybe the runner should “catch up” after coming back online. That’s easy enough to achieve if you move away from cron and track which tasks have been run in a small data store somewhere.
Yeah, I'd avoid that too.
I’ve ended up building a custom job scheduler at a couple companies I’ve worked at. It’s a fun little problem. I don’t think it’s fair to characterise the problem as a “mega” thing at all - you can build a custom task scheduler on top of a database or Kafka that’ll be way more reliable than cron in a few hundred lines of code, in just about any language.
Are those lines of code wasted? Maybe. But the trade off is that you’re the world expert in that little thing you made. It’ll be easy to connect it to your dashboards for observability, and add the exact set of features you need in your org.
I don’t think everyone should roll their own task runner. But it’s not the mega engineering problem you’re imagining. A week or two of one engineer’s time is peanuts at a company like slack.
"I wrote a CRM system in ksh script, it works great!"
Also surprised they didn’t just use an open source scheduler or product. There are a gazillion of them.
Every place I have worked cron turned into a dumpster fire.
By definition cron does not scale, it is account by account per VM/machine with no rhyme or reason.
I've been looking for software in that space, a thing that runs cronjobs, with ability to kick off adhoc runs from some UI, see what is running, and if possible parameterise the custom runs (i.e., run an ad-hoc report for a different client than usual cron does)
Plus, Jenkins has a few nice extensions to the crontab, including setting a timezone and using "H" to spread job execution load.
For jobs run by data analysts, airflow and python work great. For devops jobs, begrudgingly, Jenkins or GitHub Actions. But there's so many varieties.
I almost suffocated from all the yaml typing this, but unfortunately it’s the baseline.
Such as?
The literal evidence proves that this approach is absolutely workable at scale, so “the worst possible way” clearly doesn’t apply.
The interesting question here is “how did they make this work?” The answer to which is immensely valuable to the technical community at large. “What are they building next?”, whilst interesting, is less immediately valuable as they will be building something Slack-scale, and most orgs are not Slack or Slack-scale
The “worst possible” to me implies that it is possible.
Think about trying to solve the same problem by hurling engineers into an erupting vulcano. It is expensive, hurts the morale, causes staff retention issues, but also fundamentally does not solve the task of scheduled task running. I would not describe that as “worst possible” because it lacks the second factor by not being a possible solution.
Organisations that cobble together cron scripts for critical apps become an "an org of Slack’s size".
Companies that deploy solutions that handle all the problems with discoverability, single points of failure, failure mode options and interplanetary multi species scale when they have 0 coustomers never become "an org of Slack’s size".
You give it a schedule and an endpoint (HTTP or a Cloud Task) and then it hits that endpoint on a schedule.
I go with Scheduler->Cloud Tasks->Cloud Functions, which gives you reliability and near infinite scalability.
Very easy to reason about and full monitoring of the whole stack.
And these are just ones I personally ran into using GCP schedulers, pub/sub and functions.
See, what you're doing is re-inventing the wheel. No matter how cool your tool is - there's always work in the edge cases beyond just running the job.
We use AWS CloudWatch Events at work, and it's fine.
There is nothing about your questions that wouldn't apply to any system.
Literally all of your questions are answered in the rather well written documentation. I could go through them and answer them for you, but I don't think you'd really appreciate that.
> When designing this new, more reliable service, we decided to leverage many existing services to decrease the amount we had to build
This might explain building from scratch. Maybe the existing solutions had dependencies they didn't want to maintain and they opted for using the existing internal systems. It feels like that influenced all the rest.
I definitely don't come anywhere near Slack's scale but I've managed systems where over 3,000 cron jobs ran per day, half of which came from a cron job running every minute which usually finished in a few seconds. Some of these jobs run for X minutes too.
It's nice because there's properties you can configure for each cron job around retries and if it should be uniquely run or not. Maybe certain cron jobs should be re-tried if they fail, for others maybe it's ok to be picked up on the next interval if it fails.
Overall it's been super stable for almost 2 years which is when I started using them. Only a handful of jobs failed over this period of time and they weren't the result of Kubernetes, it was because the HTTP endpoint that was being hit from the cron job failed to respond and the cron job failure threshold was reached.
It's a good reminder that important jobs run on a schedule should be resilient to failure (saving progress, idempotent, etc.).
They all use the public curl image where I override the command in the Kubernetes cron job definition. The job container itself starts almost instantly since there's no app to boot.
If I had a case you're describing I would use the main app's image and run a specific command, in this case I'm assuming if there's not an API endpoint it would be some callable script that lives in your app's code / image.
A lifetime ago I scaled up cron jobs for a client with Gearman. Using cron to trigger jobs on the Gearman server and the pool of runners to do all the work. This proved to be so reliable they still use the system today, over 10 years later.
Crons with precisely specified time where everyone just uses whole minutes/hours are not great practice. Very unlikely you actually need such precision in a cron job and you get spiky load.
Usual approach is to set the minute to a hash of the cron config name or something, modulo 60. Hourly jobs still run hourly, but each one on random minute.
(Setting aside how fragile that setup of avoiding pod downtime sounds)
Why not smear the start time of the jobs across seconds of that minute to avoid any thundering herd problems? How much functionality relies on a script being invoked at exactly the :00 mark? And if the functionality depends on that exact timing, doesn’t it suggest something is fragile and could be redesigned to be more resilient?
As in, if you have 500 cron scripts and you think you're reaching capacity of that box, just distribute the 500 scripts in one cron tab file to two boxes with 250 each?
If one cares more about the reliability of things, you can keep tab on the cron scripts starting at their times, and if they dont, then bring the box down and start a new box with the same cron tab?
The shared feature between Temporal and those three is the workflow orchestration piece. All 3 can manage a dependency graph of jobs, handle retries, start from checkpoints, etc.
At a high level the big reason they’re different is Temporal is entirely focused on the orchestration piece, and the others are much more focused on the data piece, which comes out in a lot of the different features. Temporal has SDKs in most languages, and has a queuing system that allows you to run different workflows or even activities (tasks within a workflow) in different workers, manage concurrency, etc. You can write a parent workflow that orchestrates sub-workflows that could live in 5 other services. It’s just really composable and fits much more nicely into the critical path of your app.
Prefect is probably the closest of your list to temporal, in that it’s less opinionated than others about the workflows being “data oriented”, but it’s still only in python, and it deosn't have queueing. In short this means that your workflows are kinda supposed to run in one box running python somewhere. Temporal will let you define a 10 part workflow where two parts run on a python service running with a GPU, and the remaining parts are running in the same node.js process as your main server.
Dagster’s feature set is even more focused on data-workflows, as your workflows are meant to produce data “assets” which can be materialized/cached, etc.
They’re pretty much all designed for a data engineering team to manage many individual pipelines that are external from your application code, whereas temporal is designed to be a system that manages workflow complexity for code that (more often) runs in your application.
I'm substantially less familiar with Dagster and Prefect so can't comment as much on those.
Maybe the most data-oriented thing about Airflow is its concept of a data interval, where each DAG run is associated with some "logical date" and an interval of time that starts from the logical date (inclusive) and ends at the next logical date in the schedule (exclusive). The idea is that if you have a daily task that runs at 1 AM, then the task is expected to operate on data starting from "yesterday at 1 AM" until "today at 1 AM". But it's entirely up to the user/developer what you actually do with those logical date ranges, and you're free to ignore them entirely if you don't need them.
Dynamic task dispatch being a relatively recent feature. The fundamental design imposing lots of structure (well, kind of—you can skip lots of it, but it takes time to figure that out) to practically no benefit (and god, is the terminology dumb, made all the more so because half the stuff it names is nearly useless). “Oh yeah the scheduler just crashes or locks up while still health-checking all the time, standard practice to so restart it frequently” posted on a hundred different issues dating from yesterday to years ago (many fixed! And yet…). It’s pretty bad at passing data between tasks (see again: lots of structure, little benefit)
Operating Temporal is not that hard -- you can start with `temporal --dev` on your own box. I have a "Nomad-Temporal" Terraform Module to stand one up on Nomad. [1] Temporal has Helm Charts for Kubernetes [2]. There is also Temporal Cloud [3].
That said, there is currently a chasm between "script in cronjob" to "scheduled task in Temporal". The focus of Temporal is more "Enterprise, get your Business Processes on Temporal", not "soloist, ditch your cron".
There's certainly space for somebody to a make DAG dataflow thing or lower-code product over Temporal. Airplane.dev [4] was built on Temporal and was approaching this; acquired by AirTable.
[1] https://github.com/neomantra/terraform-nomad-temporal [2] https://github.com/temporalio/helm-charts [3] https://temporal.io/cloud [4] https://www.airplane.dev
OF COURSE everyone wants reminders at the even points, but what, do you scale up your system for every half-hour/hourly peak? Or just put in the TOS that the activation will jitter by a bit.