Show HN: Rocketry – Statement-based scheduling framework for Python
github.com
github.com
It's not spelled out, but it's apparent you just run the python file containing the app definition and leave it running in the background?
Looks very clean and pythonic.
And that's true: it's 100% Python and basically there is a main loop that checks starting conditions of tasks (and some other things) and if a task's starting condition is reached, the task is run. Tasks can be executed synchronously by setting execution as "main" or concurrently with async, threading or multiprocessing. Maybe in the future with another interpreter as well. The main loop is left running in background.
So in short, it's a Python that's constantly loop running. It sleeps defined amount of time after checking a set of tasks to lower the resource consumption but you can also create a task with execution as "main" and do sophisticated sleep like "sleep more when CPU usage is X%" or estimate the time when the next task should start from the tasks' conditions.
And thanks for the positive comment!
Rocketry plays quite nicely with FastAPI.
How does Rocketry saves execution state? Like, if it crashes and goes back up again, does it know which tasks were executed and which ones were not?
The system knows which task ran and when by extedning logging (from standard library). There is a logger called "rocketry.task" that should have a handler which can be read as well: redbird.logging.RepoHandler. An in-memory logger is created if nothing is specified. This handler abstracts simple read and write to a data store which can be an SQL database, in-memory Python list, MongoDB or CSV file.
Seems I forgot to implement a method mentioned in the docs but here's an example to specify a task log repo: https://github.com/Miksus/rocketry/issues/108#issuecomment-1...
The latest success time, starting time etc. are also stored in the tasks themselves and there is some optimization (which can be turned off) to reduce the reads in some cases. In the start-up these attributes are set in each task (if logs found).
We have a couple of hand rolled variants of this that run into all the issues this solves. Will definitely look at taking this for a spin.
Being just a lib is actually quite refreshing compared to complex behemoths like Airflow. I guess you could just use your favorite service runner (systemd, k8s, nomad, none at all...).
1. I would like to run a task each day 30 minutes before dawn so i have to compute that time at some point.
2. I run a task normally every hour but if something happens i want to run it 20 minutes after that event.
Now, nothing of that would be too hard to implement myself but these task runners pop up every so often and i would like to leverage other peoples work. Home assistant just feels a bit big for this. Do note that i currently do neither of those but i always try to evaluate these use-cases for these task-runners.
I used this myself for my father's chicken coop, it's quite easy to use.
On other side it would allow it to be fancy, like rescheduling it earlier if rain is close to the sunset
https://docs.celeryq.dev/en/stable/reference/celery.schedule...
EDIT: I think you just need to provide a nowfun that offsets datetime.now() by the desired timedelta.
If you can write the condition in English, it would seem to me you can build a custom Rocketry condition to suit it.
We've built this at https://www.inngest.com. You can run functions based off of schedules or events, with things like "when this event happens, run 20 minutes after the event". Or, "run, wait for another thing to happen, then continue".
Event driven schedulers do all the regular scheduling, but with a few benefits:
- It's reactive
- You can fan-out, so one event runs many functions
- You can store all events for debugging, replay, local testing, typing, etc.
We could plumb in an event source for #1 which indicates sunrise and sunset. Heh.
For this reason, we are using fcron[0] instead of regular cron, which allows you to specify the timezone at the start of each crontab line. If this tool supports that sort of scenario, it might be worth switching.
How does it deal with concurrently scheduled tasks and the possibility of missed tasks?
Note that this is my own opinions which probably are a bit biased. At the moment there are no built-in missed task launchers but it should be fairly easy to do such by creating a condition that checks the task run periods and whether the task did not ran the latest interval. This is not hard to do but the problem is that I haven't had time to document the time period utilities which are actually pretty extensive. I have plans and some prototypes to do pre-built a misfire condition which one can just add to any task using the OR operator.
There are 3 options for concurrent tasks: async, thread and process. Just change the execution argument of a task. Choose which suits you and remember there are pros and cons in each. All of them supports parameters etc.
It seems I did not yet implement the set_repo method even though the docs talk about this but here's one way to set a CSV repo, for example: https://github.com/Miksus/rocketry/issues/108#issuecomment-1...
Imagine I run an app with three processes A, B, C. A runs perfectly, but B fails and halts the app. If I start the app again, is it going to know that A has been already executed? Or is A going to be executed again?
There are also an option to force reading the status always from the logs. I'll provide later how's that changed but by default there is some optimization to avoid unnecessary reads from disk as often there is only one scheduler reading/writing to the log data store.
The logs are stored in memory by default but this can be changed to any data store (if you are willing to expand Red Bird). At the moment CSV, SQL and MongoDB are supported + the in-memory.
For clarity make a subfolder called "tasks" or something like that.
Then you get consolidate logging, retries and all kind of stuff for free in a battle-hardened setup and a standardized way to lookup what is enabled and what is not.
“When job a and job b are done, run job c” kind of things.
Edit: forgot the most obvious way to do dependencies… just execute A & B together as one cron job; still need something like airflow if it gets into a DAG territory
Can anyone help me with this?
https://github.com/google/python-fire
Gooey can do (CLI -> GUI):
I hope I don't lose them this time, even though I can see that I have starred one of the repos before.
The use case is wanting to have let's say a web form where a user can say they want to run a task at XYZ interval and then they can schedule and unschedule it on demand. APScheduler will pick these up without needing to restart anything.
Does your library support that? If not, is that a planned feature?
It could be the docs don't mention this. I'll need to check and add it there in case it's missing.
Or did I misunderstand?
This way you can load and unload tasks at runtime based on user input which you can optionally and independently save in your own database.
Like imagine a user wanting to control when a backup happens. You can ask them to fill out a form on your site to say "ok run this every day at 4am" and that would spawn a new job that executes at that interval and the user can also delete that and the job would be removed. There might be 100 different users each with their own individual backup jobs that are running or not running.
You can create tasks dynamically and you can create them after starting the scheduler. You can use app.session.create_task and pass "func" (Python function) for it or path and func_name if you wish to lazily load the task function (imported only when executing the task). You can also pass a command for this method as well.
And you can create a task that runs on startup (on_startup=True) and create your other tasks using this task. Use main, async or thread as execution. Then you can create other metatasks that create/modify/delete the tasks on runtime with any logic you want. For example, sync them with a database.
I'm planning on doing a proper demo about this at some point.
You can find more finer details of the repo mechanics in Red Bird's docs: https://red-bird.readthedocs.io/.
And there are methods in the session to shut down or restart the scheduler in various ways. There is also a shut condition to end the scheduling when a condition is reached.
However, this has a lot of features that Cron doesn't and which are not obvious to create yourself like create task dependencies (like "run this after that has succeeded or this has succeeded"), error management, integrating with APIs, parametrizing etc. Also if you need to run concurrently/parallel tasks, you be facing a lot of odd errors due to race conditions if you tried to do it yourself in a loop. I have even found a bug in Python's time/datetime modules while developing Rocketry. It sounds easy but I advice you don't go to the same rabbit hole as I did. Please don't, it's not good for mental health.
Of course if you need something very simple, go ahead and do it with a simple loop. Rocketry however makes easy and complex problems easy so it's still a good candidate as in case you realize your problem was more complex than you thought, it possibly has the answer or an obvious way to implement.
Compared to similar alternatives like Celery or Airflow, (I think) it is much easier to set up and more complex scheduling problems are much easier with Rocketry than with them. Of course if you are a data engineer, I suggest to use Airflow as that's the industry standard.
https://superuser.com/questions/178587/how-do-i-detach-a-pro...
Due to clock errors and accounting for thread wake wonkiness IME it's usually a good idea to have this "loop" fire in a bit of a window around when the event needs to happen (say +/- 25ms, YMMV) and then trigger the event only at the specified time. After triggering the event repeat sleeping until the next scheduled event.
Looks very clean. Is parallelism/multiprocessing (or even distribution over multiple workers) a thing you plan on making possible?
Can it support static type checking like mypy? Not only type level like str, but the time format level like “hh:mm” and “n minutes” too.
I hate when I misspell “3 minuts” and realize only after execute it. Much nice if my editor tells me.
OP, if you make one we will feature you in our blog and email newsletter (goes to around 25k devs monthly)
Background task:
I am building a service to send twitter messages daily. The limit of Twitter API is I can't send more than 1000 messages so I have to limit the call in the backend.
The closet solution I found that is I schedule a celery beat background task exactly 3 minutes and I call the Twitter API 2 times per task so daily it can send upto 960 messages in 24-hour window.
So I am still finding the solution to make this happen. Suggestion welcome.
@app.task((weekly.on("Mon") | weekly.on("Sat")) & time_of_day.after("10:00"))
def do_twice_a_week_after_ten():
Expression reads Monday or Saturday, yet its meaning is Monday and Sunday according to function name.Points in time do not actually make much sense in terms of scheduling as nothing can be run exactly at specified point. There should always be some buffer of tolerance. I think Cron has a tolerance of a minute or so and in Rocketry the tolerance is made obvious and completely customized.
For those interested more about those two types of time conditions. "time_of_..." are conditions that check whether the current time is in the specified range. The "secondly", "minutely", "hourly", "daily" etc. also check that current time is as specified but also that the task did not yet run on the interval. By combining the two you can create quite complex scheduling strategies easily.
weekly.on("Mon", "Sat")