Write your own task queue
danpalmer.me
danpalmer.me
What a load of BS. Wakatime doesn't even have tests, has been used by one person and has no error handler. It's 1k lines because it's alpha sotfware.
Celery is not the most fun to play with, but it's not big for nothing. It supports many routers AND result stores, and you can mix an match. It has several policies regarding errors, has a beat daemon, provide an API for monitoring and managing (see: https://flower.readthedocs.io/en/latest/ for UI), etc.
Also, you can use celery in a very bare bone setup using only the FS for routing and storage: https://www.distributedpython.com/2018/07/03/simple-celery-s...
I understand you may want to avoid celery. I had bad experiences with it myself. Most projects can get away with rq, or even just multiprocessing.Pool.
> Before embarking on the mission of creating a new task queue do survey the existing options, but make sure not to underestimate the hidden costs of using one, or the benefits that may come with writing one from scratch.
Writting your own task queue is like creating your own csv parser. It seems like a very simple task that is barely a for loop with a split. And then you start hitting edge cases after edge cases.
Basically you are coding an asynchronous, fault tolerant process manager including task serialization, communication protocol, priority queuing and lifecycle.
That's not something you usually want to do. The cost of using one is unlikely higher that the cost of creating, maintaining and documenting one. Just go with a simple existing task queue if you need something simple. Pypi is full of those.
As to building your own queue: if you can tolerate small loss on outages, redis + optionally a persisting replica makes things truly trivial. You can build a safe queue in a day with a small "pop into in-progress list" pipeline/script and a monitoring daemon to detect lost tasks. Monitoring daemons and dumb workers are super simple.
[1] or at least it was, a few years ago. it felt like reading an exploration of every single metaprogramming feature that Python offers, all interacting with each other.
That said, sometimes a task queue isn't the right tool for a job. At my last company, someone built something that could be called a task queue using AWS step functions. What she needed didn't exist, so she built it. Of course, now she has to respond to feature requests, bug reports, and questions about it.
Open source task queues try to solve everyone’s problems and as a result need a lot of work. By only solving your problems you may be surprised at how little code there actually is if you build on top of something like Postgres or Redis.
For a distributed build cluster that I maintain (Buildbarn, https://github.com/buildbarn/bb-remote-execution/), I also had to implement a scheduler process that would queue compilation/test actions, so that they can be executed on workers later on.
Initially I looked into using some conventional queueing system, but eventually settled on implementing my own as part of the scheduler process. So far I'm really happy with this choice, as it has allowed me to implement the following features, and more:
- In-flight deduplication of identical compilation actions. If identical actions are scheduled with different priorities, the highest priority is used.
- Multi-level scheduling fairness between groups, users in a group, builds run by the same user, etc.. The fairness cooperates well with priorities.
- Automatic removal of queued actions that are no longer associated with any running build. When the action was in-flight deduplicated, the priority may need to be lowered again.
- Stickiness, where workers prefer picking up actions that are similar to the one they ran previously, for reducing network utilisation.
- Facilities for draining workers.
Though I'm not saying it would have been impossible to achieve this with an off the shelf task queue, I'm not convinced it would have been easy. Adding new features right now only means I need to care about the actual semantics of it, as opposed to trying to figure out how to map it onto the feature set of the queueing system of choice.
Like anything else, when you need to pierce the abstraction, then it no longer serves you in its current form.
My example is I needed a task queue that did some basic rate limiting, but across workers. I wanted to be able to have 100 workers and still not hit an endpoint more than (roughly) n-times a second (or minute, or hour, or day, etc..).
There are systems that do this, but they are either complex and require complex backends and/or software installs (other platforms, etc..). Maybe they don't run in your desired environment, or maybe they do - but adding dependencies should not always be done lightly.
Creating my own task queue took a few days and then a few more days of addressing what were for us very minor bugs (duplicate execution of jobs was one).
The jobs are still technically "queued" but they don't get consumed until something opens up (but workers can work on other tasks).
queues hold jobs, the scheduler/dispatcher handles load-balancing and rate-limiting, pushing jobs down to workers only when flow-control criteria is met?
┌────────┐ ┌─────────┐
│ queueA ├──┐ ┌──┤worker1 │
└────────┘ │ │ └─────────┘
│ │
┌────────┐ ├───────────┴┐ ┌─────────┐
│ queueB ├──┤dispatcher ├─┤worker2 │
└────────┘ ├───────────┬┘ └─────────┘
│ │
┌────────┐ │ │ ┌─────────┐
│ queueC ├──┘ └──┤workerN │
└────────┘ └─────────┘My system supports as many dispatchers as you want, which is a good addition but makes the logic more complex as you have to be careful with locks so you don't schedule jobs many times, for example.
My system also implements retry logic, and a bunch of other stuff, but that isn't absolutely required either.
Now Elixir with Genservers that can handle proper concurrency and back pressure is trivial.
One of the best is Oban, which uses PostgreSQL for the queue storage. No more celery, Redis or other external dependencies needed.
Heck if you don’t need persistence you can use Que which keeps it in ram.
That being said I would never write a task queue any longer. Too many good options out there now for every major language.
We’re not Google.
1. Queue jobs in a database table so you get transaction guarantees.
2. Use a prebuilt system that works with any number of existing alternatives (or one that works with the database too).
I can’t imagine building a queue that meets any slice of my requirements to be able to trust it without heavily leveraging SQL.
Also, am I crazy or do the celery docs not even clarify their delivery semantics? Isn't that table stakes for a queueing system? As best as I could tell you can get close to "at least once" with
acks_late=True, task_reject_on_worker_lost=True
but not in cases like a worker hanging indefinitely without being explicitly killed.Then, add a "sweeper" that checks if the task actually happened and if it did not happen, requeue the job.
To ensure proper behaviour that the same job is not being worked on twice, add some locking in there and you have a really stable way of scheduling jobs.
Could this have been done INSIDE celery? Sure, but I would say that this is not the case that it should optimize for.
> We have a task queue…
> Is it Celery?
> No
> oh thank god. How do I use it?
> Put the @task decorator on a function and call it. See the args to that decorator if you want slightly more control.
The networking on our redis instance was maxed out all the time and we had a secret soft-cap on the number of workers that could be active.
The documentation even says that as-is this "gossiping" doesn't even do anything!
Unsure if it's worth it for us to write our own, but I felt inspired to share my Celery horror-story.
Because I’ve built my own “task queue” on top of “SELECT * FROM tasks WHERE done = 0 ORDER BY create_date LIMIT 1;” and have also seen others do it. Just don’t do it like this. You’ll get spikes of tasks and then your table will be filled faster than your worker can finish the tasks. Just use something else.
I wrote my own in half a day, then spent a few days total tinkering with it too, so far so good.
Write your own, and even if you give up halfway through and use the popular library after all, you'll have learned a lot about how stuff works, and can make better decisions about how to use the library.