PinLater: An asynchronous job execution system
engineering.pinterest.com
engineering.pinterest.com
Rabbit's advertised throughput doesn't reach 100,000/sec, but you should be able to shard your way around that. Or use Apache Kafka, which has the sharding built in.
http://www.slideshare.net/ToddPalino/enterprise-kafka-kafka-...
Celery has flexibility for message backends, but prefers RabbitMQ which comes with its own operations overhead, and does not have real horizontal scaling built in (it has failover techniques). RabbitMQ is great for mission critical messaging that survive system restarts.
My biggest worry is relying on Redis or MySQL. What happens if workers stop consuming tasks, wouldn’t that balloon the redis memory consumption pretty quickly? I suppose memory consumption is a factor of load and task bytesize. MySQL on the other hand feels like the wrong tool for this job.
You can write simple workers for Celery in other languages too, and there is one in active development for Java. But then Redis is not really a good choice, AMQP is more convenient for interoperability.
Have no idea if pinterest evaluated Celery, because they have not been in contact. If they did there are several pitfalls they could have gotten wrong at this scale, but I'm pretty confident it would have been more than suitable. It could have been a benefit for all of us and chances are they have underestimated the effort needed to maintain a custom solution.
* Yes, we did evaluate a few open source solutions before deciding to build in-house. We couldn't at the time find any solution that met all the requirements that we had: scale, scheduling jobs ahead of time, transactional ACKs, priority queues, configurable retry policies, great manageability (e.g. ability to rate limit a queue online, inspect running jobs, review failed job stack traces). [update: also support for clients/jobs in multiple languages: Python, Java, Go, C++]
* Note that we were building a v2 system. Not improving substantially on the existing system on multiple dimensions would not quite make it worth the cost of doing large scale migration of tasks. We had to meet all the requirements, could not compromise on missing a few.
* Sorry I didn't mention explicitly in the post, but we do plan to open source PinLater, hopefully later this year.
Happy to answer any other questions or specifics!
Basically these guys 'invented' the queue.
Also, the problem of acknowledging success without blocking can be solved without queues - by programming in async environments like NodeJS
kind of. It only re-queues if you drop the socket, I believe. Just sending an error status will not have the message re-queued sadly. I've looked into this, but everyone seems to be doing that manually.
We have been using for years to process 1+ million/jobs per day without a single issue.
"While traditional messaging systems tend to use one of the models described above ("broker" model in most cases) ØMQ is more of a framework that allows you to use any of the models or even combine different models to get the optimal performance/functionality/reliability ratio."
Therefore, I don't really understand the purpose of this blog entry. Yay, they built something. What's the use, if it's not published and usable for others, too?