HNHacker News
TopNewBestAskShowJobs

Miksus

118 karma · joined August 27, 2022

Creator of Rocketry, developer in a trading room
submissionscomments
Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Sorry, I'm in quite a hurry (so sorry for the language and lack of ellaboration).

You can create tasks dynamically and you can create them after starting the scheduler. You can use app.session.create_task and pass "func" (Python function) for it or path and func_name if you wish to lazily load the task function (imported only when executing the task). You can also pass a command for this method as well.

And you can create a task that runs on startup (on_startup=True) and create your other tasks using this task. Use main, async or thread as execution. Then you can create other metatasks that create/modify/delete the tasks on runtime with any logic you want. For example, sync them with a database.

I'm planning on doing a proper demo about this at some point.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
That's a great idea actually. Thanks, I think that should be pretty easy to do!
Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Do you mean with "dynamically schedule and unschedule" that you sort of manually (or using another task) stop running a task in its specified interval (or condition)? It does support this, there is an argument "disabled" in the tasks that can be set True and then the task won't be run unless explicitly forced to run (calling run method of a task). The task can be enabled by setting it back to False. This can be done in runtime in another task using main, thread or async execution.

It could be the docs don't mention this. I'll need to check and add it there in case it's missing.

Or did I misunderstand?

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Thanks a lot, nice to hear!

The system knows which task ran and when by extedning logging (from standard library). There is a logger called "rocketry.task" that should have a handler which can be read as well: redbird.logging.RepoHandler. An in-memory logger is created if nothing is specified. This handler abstracts simple read and write to a data store which can be an SQL database, in-memory Python list, MongoDB or CSV file.

Seems I forgot to implement a method mentioned in the docs but here's an example to specify a task log repo: https://github.com/Miksus/rocketry/issues/108#issuecomment-1...

The latest success time, starting time etc. are also stored in the tasks themselves and there is some optimization (which can be turned off) to reduce the reads in some cases. In the start-up these attributes are set in each task (if logs found).

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
If you want an API (or UI), just clone this and modify it as you need: https://github.com/Miksus/rocketry-with-fastapi. I also wrote an article to Medium how it works with FastAPI: https://itnext.io/scheduler-with-an-api-rocketry-fastapi-a0f...

Rocketry plays quite nicely with FastAPI.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
You can do without a database backend but of course then the task logs are not kept in case of restart. Currently you can use any SQL database that SQLAlchemy supports, MongoDB or CSV files, or any other if you wish to extend Red Bird. It uses Red Bird (another project of mine) to abstract the data store: https://red-bird.readthedocs.io/en/latest/. And it just extends the logging library for reading task logs.

It seems I did not yet implement the set_repo method even though the docs talk about this but here's one way to set a CSV repo, for example: https://github.com/Miksus/rocketry/issues/108#issuecomment-1...

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Putting it mildly, this is nothing more than a sophisticated Python while loop. And it's not as performance friendly as Cron due to that Rocketry runs on Python. You need to be able to leave Python program running in order to use Rocketry. As bad as that sounds, it's not really a problem with modern machines though. Have run this on Raspberry and with a machine with even poorer specs.

However, this has a lot of features that Cron doesn't and which are not obvious to create yourself like create task dependencies (like "run this after that has succeeded or this has succeeded"), error management, integrating with APIs, parametrizing etc. Also if you need to run concurrently/parallel tasks, you be facing a lot of odd errors due to race conditions if you tried to do it yourself in a loop. I have even found a bug in Python's time/datetime modules while developing Rocketry. It sounds easy but I advice you don't go to the same rabbit hole as I did. Please don't, it's not good for mental health.

Of course if you need something very simple, go ahead and do it with a simple loop. Rocketry however makes easy and complex problems easy so it's still a good candidate as in case you realize your problem was more complex than you thought, it possibly has the answer or an obvious way to implement.

Compared to similar alternatives like Celery or Airflow, (I think) it is much easier to set up and more complex scheduling problems are much easier with Rocketry than with them. Of course if you are a data engineer, I suggest to use Airflow as that's the industry standard.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Yep, exactly. Rocketry works on conditions which are either true or false thus you need to give it a time range.

Points in time do not actually make much sense in terms of scheduling as nothing can be run exactly at specified point. There should always be some buffer of tolerance. I think Cron has a tolerance of a minute or so and in Rocketry the tolerance is made obvious and completely customized.

For those interested more about those two types of time conditions. "time_of_..." are conditions that check whether the current time is in the specified range. The "secondly", "minutely", "hourly", "daily" etc. also check that current time is as specified but also that the task did not yet run on the interval. By combining the two you can create quite complex scheduling strategies easily.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Of course I have this running though it's still running older version (been too busy with developing this). It has been running over half a year for my scrapers without a single interruption even though the machine has the worst specs available. I have tested this with Linux/Unix and Windows at least. Of course, I have gotten message from various people saying they are using it. Some have said they migrated from Celery or other alternatives as they found Rocketry more suitable for their needs.

And that's true: it's 100% Python and basically there is a main loop that checks starting conditions of tasks (and some other things) and if a task's starting condition is reached, the task is run. Tasks can be executed synchronously by setting execution as "main" or concurrently with async, threading or multiprocessing. Maybe in the future with another interpreter as well. The main loop is left running in background.

So in short, it's a Python that's constantly loop running. It sleeps defined amount of time after checking a set of tasks to lower the resource consumption but you can also create a task with execution as "main" and do sophisticated sleep like "sleep more when CPU usage is X%" or estimate the time when the next task should start from the tasks' conditions.

And thanks for the positive comment!

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
Yep it does, the main process and thread is responsible of communicating with the logs (see the other comment in which I explained the logging mechanism). If you run a task in subprocess, the logs are relayed via queue to the main process and the main process logs it to avoid conflicts.

There are also an option to force reading the status always from the logs. I'll provide later how's that changed but by default there is some optimization to avoid unnecessary reads from disk as often there is only one scheduler reading/writing to the log data store.

The logs are stored in memory by default but this can be changed to any data store (if you are willing to expand Red Bird). At the moment CSV, SQL and MongoDB are supported + the in-memory.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
There is a repository mechanism to store the logs. The task logger is simply an extension of logging library. Seems my docs are slightly off on setting up the CSV repo but you can just add a RepoHandler (from redbird) to the logger called rocketry.task. At the moment there are MemoryRepo, CsvFileRepo, SQLRepo and MongoRepo.

You can find more finer details of the repo mechanics in Red Bird's docs: https://red-bird.readthedocs.io/.

And there are methods in the session to shut down or restart the scheduler in various ways. There is also a shut condition to end the scheduling when a condition is reached.

Miksus··on Show HN: Rocketry – Statement-based scheduling framework for Python
I wrote my own ideas of how it compares to APScheduler (and other alternatives) here: https://rocketry.readthedocs.io/en/stable/rocketry_vs_altern...

Note that this is my own opinions which probably are a bit biased. At the moment there are no built-in missed task launchers but it should be fairly easy to do such by creating a condition that checks the task run periods and whether the task did not ran the latest interval. This is not hard to do but the problem is that I haven't had time to document the time period utilities which are actually pretty extensive. I have plans and some prototypes to do pre-built a misfire condition which one can just add to any task using the OR operator.

There are 3 options for concurrent tasks: async, thread and process. Just change the execution argument of a task. Choose which suits you and remember there are pros and cons in each. All of them supports parameters etc.