A Fast, Simple, Queue Built on MongoDB
blog.attachments.me
blog.attachments.me
* What happens when job's crash, fail, or get stuck?
* What about when the machines running the workers die or are isolated from the network?
* What happens if the backing store dies?
* How do you start, stop, restart, grow, and shrink your worker pool?
[1] http://ask.github.com/celery/getting-started/introduction.ht...
* We are a Ruby/JavaScript/Python shop, so a library like Celery is not an optimal solution.
* MongoDB is a core part of our infrastructure if it dies and is unavailable, we have bigger fish to fry.
* Every piece of architecture in a stack is conceptual and maintenance overhead. I am certainly not averse to learning new technologies, but for the time being using MongoDB for this makes sense.
I should underscore that most of the actual queue implementation is simply taking advantage of the behaviour of capped collections in MongoDB. Karait is a convenient API on top of this.
IMHO, you're making a pretty big assumption here anyway with the fixed queue depth. The idea with a message queue is that it delegates the responsibility of ensuring message delivery to another component, so that enqueues always work, and that messages always get delivered. Perhaps it's ok in your use case that some tasks fail to get delivered, but I'd prefer not to voluntarily give that up for some reason that I may have missed from the original post. In the linked .NET example, they used capped collections because it was a chatty message bus, not a worker queue, and guaranteed message delivery wasn't important. I've nearly filled the drives up on a RabbitMQ box with messages and it was still able to insert and pop them off at thousands per second, with full durability and persistence.
Celery can read and write JSON to a multitude of backends. Last time I checked, both Ruby and JavaScript have pretty solid JSON support. DaemonKit on Ruby supports building a robust worker pool over AMQP.
Every wheel reinvented is also conceptual and maintenance overhead, and I'd argue that, most of the time, is probably far worse than writing tools from scratch. There's a highly likelihood reinvention will introduce a bug or poorly-handled "edge" case in a from-scratch tool than if an existing, mature tool is used. When "bad things happen," a mature tool will give "a way out" that a custom tool probably won't have. Much like the honey badger, when scale and disaster strikes, it simply does not give a shit about how minimalistic your stack is. Good tooling can make short work of navigating out of these messes.
In my use-case I am simply passing a message to a worker to indicate that it it must terminate. There are no adverse consequences if the 'job' fails.
I am most certainly not advocating that Karait solves all the world's queueing problems. It solves a specific problem I have, and I am confident in this assertion.
BTW, do you get some sort of fiduciary kick-back from the makers Celery ;)
Having said that, this is a good temporary solution for me to shutdown the crawlers. And I see absolutely no harm in abstracting the use of capped collections into an easy to use library.
Is it as well developed as a mature AMQP queueing solution? Of course not, I don't claim that it is. Having said that, I trust MongoDB in my stack as a means to communicate between processes, and am confident using it for the use-cases that I am. As long as I'm doing this, I see no harm in open-sourcing the helper libraries I create for doing so.
Am I a huge advocate of MongoDB? No, but if you happen to have it in your stack, as I do, It's handy to be able to use it in a similar manner to the way that Resque leverages Redis.
We get to look at lots of people's Rails apps, and it appears that Resque is emerging as the de facto answer to this problem. A nice attribute of Resque is that for all the machinery Redis provides, it also solves the "memcache" problem and is probably a better utility player than $X-MQ.
If you have an app that already wants $X-MQ, Celery sounds very sensible. And having apps that want $X-MQ is a good thing, too.
And, not reinventing this wheel also makes sense.
How is RabbitMQ a lot of machinery? Compared to Redis, AMQP and Rabbit have the complexity of a bologna sandwich. It's actually a quite simple system, and very easy to install. There is almost always no configuration necessary. Developers can install it in one command with MacPorts or Homebrew or apt or yum and don't even have to create users or "vhosts" to get started. Create a named exchange, subscribe to it, you're now AMQPing. It's well-packaged and well-documented. It's also so light on it's feet that it rarely actually needs a dedicated box; just stick it on your database server to get started.
I think Resque is a good solution for Rails, for sure. I've built stuff around it, and it works well, and definitely battle tested. It certainly beats the existing AMQP-based solutions. The idea that your application code probably sucks and/or is in a constant state of flux, so handle failure gracefully is rather excellent. The web interface is particularly nice and useful. While I don't necessarily agree 100% with Redis as the backing store, the actual mechanics of the worker queue implementation (plus tools like resque-pool) work so well that for most projects, it's perfectly fine.
I actually had a problem specific to using Redis materialize when it ran out of memory while queueing a large batch of per-user-generated weekly emails. A cron job would spin up a script that would cache a large amount of database data in RAM, using it to generate per-user content. The rate at which email could be sent became a major bottleneck, and we had Resque, so we just enqueue'd the entire email and had a small job that connected to the email server and send them. This worked reasonably well and for many months, the high watermark for RAM usage was well below what the box could hold.
Unexpectedly, the amount of content contained within these emails quadrupled. Redis ate through all the RAM on the box and it started going to disk using Redis' VM mechanism. This was configured to "theoretically" prevent an unexpected OOM from making the queues unavailable. The unfortunate part is that the Resque stores the entire queue in a single key, and when the key gets swapped onto disk, Redis must read this entire key (and therefore the entire queue) into RAM before it's able to read any data from it.
This wasn't Resque's fault really. It was my fault. I did not take the time to understand how Resque and Redis worked well enough to make a truly informed decision. I did not test the OOM condition. I should have tried to fill up the queue on a test box with a billion items just to see what happened. All lessons learned.
[1] http://www.rabbitmq.com/blog/2011/01/20/rabbitmq-backing-sto...
Celery has multiple backends, including Redis - we've been using the Redis backend quite happily for lanyrd.com and it's worked like a dream.
Are you guys Python? Is Celery+Redis the Pythonic Resque?
I'm saying that as the author of Delayed::Job so i'm pretty much responsible for those shenanigans.
If you look at Twitter's Kestrel, it was implemented with a similar mechanism, albeit leveraging some capabilities that were intrinsic to memcached.
I built a little queuing prototype using long polling and it works very well. If you build it using an async server, you can run many clients while only using a few threads.
There's still the question of how to block and/or wake on an enqueue event to MongoDB vs. polling. The best approach may be to implement queueing using your REST API, leveraging all of the advantages of the async server and use MongoDB for a backing store or journal. This is similar to how Kestrel does it.
My prototype was very responsive, although I only got about 500 messages per second of throughput on my MBA.
I was thinking of doing a proof of concept in Node, that would at least behave like it's getting pushed messages rather than pulling ... But, I couldn't think of a way other than repeated setTimeouts.
Any thoughts?