Three quick tips from two years with Celery
library.launchkit.io
library.launchkit.io
My life got better when I stopped using Celery and instead started using Redis or RabbitMQ directly. It's not that hard.
Once deployed in production I spent countless evenings debugging why the tasks weren't running, why upon timeouts tasks weren't terminated, and why the task queue was getting bigger and bigger. 50% were configuration problems, due to the _bad_ and confusing documentation, 50% were due to introducing a sheer amount of complexity with Celery that I often spent my time reading through the source code to understand what was going on. And Python, with all its magic methods and abstractions, made that very hard.
One day I just rewrote everything with three Django management commands: two for the background tasks, one for the schedulation (with multiprocessing). I haven't heard from the client since.
However, when it starts hiding exceptions within your tasks if you use JSON for messages or if you are having problems with it doing things you were not expecting... Like randomly not passing information between dependant tasks... Gevent and celery I have found don't play well together even if you disable Gevent for the worker.
Maybe you eventually realise that writing your own task for rabbitmq and your own workers is simpler than relying on someone elses complex code.
Celery will 100% do what you want but it's definitely been a steeper learning curve than it should be and has lots of gotchas and many hours of hunting around to deal with obscure issues.
I do wonder if there is a reasonable heuristic, apart from experience, to learn when a piece of software may be too complex for at least me to use well.
Curious how our use patterns are different that causes what you're seeing.
Is there any truth to this ?
The language for building workflows is pretty half-baked, and there have been times when I've wondered whether RQ or raw message passing would be better in my situation. But it's better than any of the alternatives that I've had direct experience with.
I'm not sure I'd look to something that explicitly uses Redis as an alternative unless I could be sure the queue would be reasonably small.
Which, if the jobs are not going to be overwhelmed with backed up jobs, maybe that's fine.
Also, maybe we were doing it wrong.
> This option comes with a coordination penalty, but results in a much more predictable behavior
In general, I'd much rather the greater throughput over predictable behavior. In general, I'm not watching my queue and having it be predictable isn't a very high priority.
Assuming most Celery task queue work involves network requests that can be unpredictable, I think Ofair is a more sensible/understandable default.