Celery 4.0
docs.celeryproject.org
docs.celeryproject.org
I imagine as native async tooling improves in py3x (async/await, aiohttp, and other tools) use of celery to do trivially concurrent things will decrease and Celery's usage will focus on more complex workflows (chords, fanouts, map/reduce).
Looks like many concerns we're tackled here (thanks Celery team) and I'm looking forward to playing around with this release.
I understand there's a genuine use-case for Celery but like many technologies people are told "if you need a task queue use this" when there are much simpler solutions that are more than good enough for most.
Disclaimer: I'm a contributor
Its useful to know Celery, and gets used in a proper context in my current work, so I guess learning it wasn't a waste.
Joking aside, we use Celery in production for asynchronous number crunching. We build a web UI in Django that raises Celery tasks that run long number crunching jobs. When the number crunching is done we look at the report through our Django web UI.
It works well enough, though we wish there were better built-in task management. E.g. A built-in API to reshuffle tasks (e.g. for I/O resource balancing) would be nice. celery purge also seems too drastic in purging every task in every queue. Or maybe we're just doing it wrong.
It's kind of like jQuery's ajax. Of course you can figure out how to make a simple replacement. But then you have to manage the 100 different edge cases for when your code doesn't go down the happy path.
Much easier to write a simple task queue that's good enough than a replacement for $.ajax, though...
The best thing about celery is that you don't have to use any of the advanced features. If you need the basics, stick to the basics.
On the flipside, if and when you do need more than the basics, your system will grow with you. No need to hack up cron monstrosities as you grow beyond a single server. Though, cron vs celery is kind of an apples to... celery comparison.
Yeah - everyone should ideally be comfortable with all this stuff but I try and keep anything that's more complex than a pip install in a virtualenv to an absolute minimum.
If you are on AWS, you could just use something like their email service to fire off the request, then you don't need a queue at all.
You could add your emails to a db table and have a cron job consume them.
But for sending an occasional email? I've never really had a problem with just connecting to the SMTP server in the request/response cycle. If it takes more than a second to send then something is seriously wrong. You could use Mailgun or similar services which have their own queue for handling bulk sends and further reduce the likelihood of a problematic blocking of the web-server.
Maybe I will try something simpler like http://stackoverflow.com/a/4447147
This is a small detail, but I'm glad that one more big Python project is dropping 2.x support. The 2/3 split is not good for the community.
Edit: Apparently they curtailed some features for simplicity as well.
I have merged many features, like broker transports, result backends, etc, and while the initial contribution was great, it ends up being unmaintained with issues that nobody fixes.
If there's any feature that you really want back, chances are the problems with that feature are not super difficult to fix, so please reach out!
In the end I shouldn’t have been, it was a pleasant surprise. It's great 4.0 has come out.
Almost every Django developer will touch celery at some time, if your organisation can support development please do, it's not just an important piece of software, but well put together too.
I'm not a heavy Python user, and I've never heard this before. It sounds... less than good.
requests is used so widely that I'm surprised to see this raised as an issue.
Some time ago I built a little celery addon library as a sort of experimental way to solve the problem of having dynamic celery beat scheduled tasks. I never ended up implementing anywhere for a few different reasons: https://github.com/fuhrysteve/CeleryStore
I really like the concept behind the old djcelery project. But I don't use django much these days, and I'd like for it to be more compatible with tools I've become more familiar with (sqlalchemy / etc).
Do you have any advice for how to approach this? I know 4.0 introduced some new abilities to add beat entries via API.
[1] CELERY_SEND_TASK_ERROR_EMAILS config removed http://docs.celeryproject.org/en/latest/whatsnew-4.0.html#fe...
[2] https://gist.github.com/alanhamlett/dc8cdd4721ea63053f14#fil...
For AWS they have SQS which is awesome. Now if someone would write a Google cloud pub/sub, all would be well in the world.
Would love to see how they stack up after this release.
The reason we switched to Celery was the volume of messages we were handling. Since Rq relies on Redis, all your messages need to fit in memory. While Rq was great and simple to setup at start, as we grew we were consistently dealing with Rq breaking because Redis was full and stopped accepting any write operations.
We moved to Celery because it could use RabbitMQ as a broker. RabbitMQ offloads most messages to the disk which has nicely taken care of the memory limitation issues.
With Rq we would get stuck after 10K messages (our messages included images so individual message size was large). With RabbitMQ I've seen the queue grow to about 120K without so much as a single hiccup.
You can do such things as passing a database unique key, GUID, or file-path to the raw data on disk. Obviously, you will also need to engineer around that if you've got a distributed system. The tangential benefit is that you're not using a "messaging" queue or system for persisting or semi-persisting your image data. That's a big no-no as such systems are transient in nature and that often doesn't align with binary or image processing.
Base64 is for when you want to pass around binary in a text-based format. E.g. XML or JSON. But do keep in mind that because of the encoding format, converting binary data to base64 does increase the payload size by about 15% or so.
Unfortunatelly version 4.0.0 has too many bugs for now. I will need to check back in few months. Better tests and better mocks would prevent the issues.
Honestly this is a complex problem area and I think the Celery developers have done an excellent job of making it pretty trivial to get up and running while providing lots of flexibility for more advanced users! Not an easy feat.
Are there bugs? Of course - but I've never come up against one that I can't work around. Is that annoying? Sometimes but that's software development.