Show HN: Your Own Task Queue for Python
github.com
github.com
Don't get me wrong, for basic stuff celery is great, albeit with a bit of a learning curve for the beginner. (And who knows, maybe I'm just missing something when it comes to more complex applications.) I also think the devs have done a great job and are super responsive. I think they also made a wise choice to chop features a little while back.
What I would really love to see (create?) is a cut-down python task queue which deals solely and specifically with some kind of AMQP (or similar) broker. Also with a task scheduler with a UI, that is something I'm loathed to lose.
Maybe this seems ridiculous, I get the impression that everyone is talking about microservices, ZeroMQ, service discovery, billions of messages per picosecond etc. I'm talking about medium size projects where a maybe a dozen apps just need to coordinate a bit. I feel like this middle ground lacks solutions.
Sorry for the ranty comment. I suspect I'm just trying to talk myself into creating this at some point.
[1] https://www.rabbitmq.com/tutorials/tutorial-five-python.html
I wanted to create a task for my web app to run every 15 minutes and I didn't realize how complicated it was going to be. I'm using AWS beanstalk and I wanted the task to run on 0,15,30,45. The "simplest" way to do that was to add a worker box and create a cron job (it uses AWS SQS as an intermediary). I want to make sending email tasks as well but that means I'm going to have to move deeper into SQS integration and I don't want to be stuck with AWS. My only other option seems to be setting up and maintaining rabbitmq/celery and that looks like a nightmare. :/
The hardest part seems to be answering the "Why isn't X receiving messages from Y?" question. Although it didn't help that I got '*' and '#' switched around in my head (seriously, '#' is the wildcard?!)
I'm also running little celery instances on each web worker, using Django as the backend. No SQS/Redis/RabbitMQ, no huge AWS lock in, no having to set up a backend worker layer. In general it's been fine for a year. You don't get good visibility into processing queue length etc but we usually have four bone-idle web workers so there is no queue length. If that might be a fit for your situation I could bang together a blog post or something.
Cron would probably be a better bet then.
I love celery but I'm not a fan of celery beat.
This bit of code in worker.py makes me uncomfortable though:
# Deserialize the task
d_fun, d_args = dill.loads(data)
# Run the task
d_fun(*d_args)
Basically the function (as in the Python object that holds the function's code, etc) to run gets pulled out of redis along with its arguments, then it gets executed.If the entire system is trusted, this is a really cool way to let workers run code without needing to deploy it to them. Likewise, you wouldn't even need to restart the workers for a new version of the function to run - just start feeding the new code into the queue and the workers will run it. The more I think about this, the cooler it seems.
On the other hand, from the worker's perspective, this is basically a wide-open "feed me any code and I'll run it" thing.
A few years back we were using a service called PiCloud that effectively allowed you to do that. It went a whole lever deeper and actually bundled and shipped your local code to the remote runner so you could actually run a single function from your codebase remotely. It was a one of the best products I've used but a) their service was unreliable and b) the got acqui-hired (by Dropbox, I think).
It was something simple like:
import cloud
def my_long_fun():
.....
cloud.run(my_long_fun)
and all the deps and everything would be shipped. Next time you called a function, if it already had everything it would be able to run remotely without shipping again. import cloudpickle
pkl = cloudpickle.dumps(my_long_fun)
# Ship it somewhere
new_fun = cloudpickle.loads(pkl)
new_fun(...)
https://github.com/cloudpipe/cloudpickleIt was so wonderful. And I was so disappointed that Multyvac (http://www.multyvac.com/) dropped the ball. I thought I'd be able to continue using the service ...
BTW, Multyvac has one of the more hilarious "trusted by" sections I've seen on the internet.
I spent a couple of weeks replacing it with celery, and then later RQ. It was the correct call, but you never know with these things. As far as it goes it was a pretty cheap lesson in not relying on external providers. Having said that, I've spent a fortune on EC2 since because I can't have micro billing like I could with picloud!
Not that Celery is the only way to go. This is good work - I'm just interested in the why.
I'm sure they exist in theory, but it's overkill in every situation that I've seen it deployed.
also, rabbitmq has an amazing admin interface, monitoring is really easy (most things you use for monitoring support it) and the whole thing has really sane defaults (imho).
all in all, it's one of my favorites pieces of software ever!
(btw, i remember people having problems with celery because of they way they imported the tasks. if you always do full imports (as in, from a.b.c import tasks instead of from . import tasks) you shouldn't have any issues).
[1] https://www.rabbitmq.com/getstarted.html
PS. I'll stop evangelising soon, I promise.
At the time (it was a couple of years back) there was no way to stop it from taking 2 tasks at a time from the queue. We have a smaller number of long running tasks and it made it very difficult to scale out the work. Also, the monitoring story was pretty rubbish and debugging was hard.
Now we use RQ, which is mostly pretty good but we've been slowly adjusting parts of it's behaviour to make it more suitable for us. Recently, we did some work around how it does it's TTL so we could more quickly see where workers had terminated abnormally.
We've also changed it so a given worker instance will become increasingly unlikely to take on new work as the machine it's running on (AWS) nears the end of the hour since it was started. As the work load lowers, we can remove workers from the system in the most cost effective way.
I was hoping someone could kindly explain what a task queue is used for. If I understand it correctly, a task queue would be used for running automation code that is written during the development of a python project? But that doesn't sound right.
Can someone please tell me what a "task queue for python" means in context
Imagine you are GitHub and a user comments on an issue that has 1000 people participating. As soon as the user makes a comment you want to send an email to every user that there are new comments.
The straight forward way is to do that in the web application when the comment is posted, but that takes a lot of time and leaves the commenter staring at a loading screen until everyone is notified. (There are other reasons as well)
With a message/task queue you would create 1000 messages (each saying "send e-mail about this issue to this user"). Then you have worker processes that in a loop look at one message at a time and sends an e-email. When the message is sent the worker marks (acks) that message as handled and it is removed from the message/task queue.
What's your experience been using this and Redis for your queue backend?
https://blog.2ndquadrant.com/what-is-select-skip-locked-for-...
assuming you are happy to serialise things as JSON or pickle then this is pretty easy to integrate with python
The various options all have pros and cons, and fit different problems better or worse. It's worth trying a few different ones out to see what's best for you.
A minimal celery deployment is 1 or two configuration options. Celery does have a metric ton of moving parts but distributed systems are hard and celery can do a lot of things.
This isn't the first time I have seen this sentiment. Perhaps there is a documentation problem to be solved and/or a basic API that could be provided by Celery?
Quick questions:
Can the worker push tasks back onto the queue e.g. using r.lpush('tasks',mytask)? This would be useful in a try-catch where the worker fails its task and needs to post it back.
Can the worker send its results back to Redis, it is just r.lpush('finished',...)?
Very true :-)
And if you just want a simple broker that does the right thing, take a look at beanstalkd - I built a system to mimic GAE's queues awhile back and beanstalkd was very nice -- designed in spirit of memcached where it has a very simple api and aims to do one thing well.
Edit: no closures, I think