Rescuing Resque: Let's do this
blog.steveklabnik.com
blog.steveklabnik.com
Our solution was to build yet another queueing and worker system. For queues Beanstalkd is way faster and way more reliable than Redis. Instead of EventMachine I built the workers with Ruby threads and making a retry system on top of it wasn't really a problem.
Our error percentage dropped from around 10% to 0.0001% on average. Maybe some day I might release this as a gem, if world really needs yet another queueing system.
I respect your opinion on different messaging solutions (Resque VS Beanstalkd) however since we spend a lot of efforts to make sure Redis is very reliable I wonder what's the reliability problem you experienced with Redis. Thanks!
I'm not sure a whole lot can be done, other than resurrecting the old debate over something like diskstore. And while EC2 is memory-constrained, dealing with this situation isn't a matter of just adding more memory, because that will be filled up as well. redis would either need to be able to spill over onto disk or the processing model would need to be changed. We ended up doing the latter, eschewing stacktraces for retried jobs, adding monitoring to disable queues that have a very high failure rate, etc.
Redis is very good for different data structures, like hashes or lists, but for queuing I really prefer Beanstalkd over anything.
As far as perf goes, redis queuing is just a linked list. It is ridiculously fast, I would really like to see numbers backing up your claims.
If you are going to make accusations like this you should at the very least understand why stuff broke.
I prefer simple solutions and I know it was our fault we got a data loss with Redis. Still Beanstalkd queues have been causing less pain than Resque and Redis together. Integrated delays, tubes and deadlines are really nice.
Altogether the reliability comes from the deadlines. If my worker gets stuck (well, it's Ruby), we don't lose the work item. Doing pop on Redis will delete the item from the queue, but in Beanstalkd you reserve the item and delete it when you're done.
Before pushing a job to beanstalk we encode the entire job a 'current-jobs' hashmap in redis, our worker transfers it to a 'busy-jobs' hashmap before processing, and then deletes it if successful or transfers back to current-jobs if failed gracefully.
This way we get the full useful features of Beanstalk (tubes, delays, etc) while knowing that our jobs are always sitting in Redis too if the beanstalk server dies (its persistence features are lacking) and our commit function is atomic (Redis MULTI/EXEC).
We're using Redis for our big document processing-and-indexing pipeline at Cue, and it's great software, but it's not a ready-made queueing system. All of the features I mentioned above are things that we've had to build ourselves. Redis is more like a general-purpose building block for all kinds of data systems.
Our queue items are plain id values, which trigger a set of actions from the database to the internet. If there's a database failure or the process itself crashes it is very nice to know our reserved work items will not be gone but released back to other workers to process.
> And then there's all the front-end goodness for debugging: you could easily want things like graphs for flow and queue lengths.
I have lots of graphs from the system in Graphite. Works perfectly.
> And how about master-slave replication, failover, sharding, and so on?
Sharding is easy to do, just specify an array of servers in the clients and the clients will shard.
Replication is a bit different problem. The server will write a binlog file to the disk, which is then backed up. We have beanstalkd servers waiting in another machine pointing to the same binlogs. On an error situation we just switch them on and set our routes differently.
Yeah, the biggest problem with Beanstalkd is the missing replication, but we can live without it.
> We're using Redis for our big document processing-and-indexing pipeline at Cue, and it's great software, but it's not a ready-made queueing system.
We're not doing so heavy processing, but we're relying on many third party services, which can fail randomly. The retrying system is a must and of course we cannot miss our work items so often, so the deadlines help.
And of course the tubes rock. Now I can build a separate tube for each retry, I can set different priorities for the tubes (so the more fail prone tubes won't block all workers), I can set a deadline for a job to finish and if the worker dies in the middle the job will be returned for the other workers. And best of all, I can monitor all tubes nicely with Graphite.
1. When doing something, increment a counter in Statsd, which will aggregate the values to Graphite. Works well for custom stuff.
2. Write a rake task to poll Beanstalkd with `stats` command, reduce the data and send it straight to Graphite.
I'm really planning to open source the queuing system we're using because it's so simple. It might need a prettier admin view and some nice way of plugging in more monitoring systems before I put it to public.
EDIT: It seemed familiar, so I dug up an old thread of yours [1], where we discussed a similar thing. At that time, it seemed you narrowed down your problems to Eventmachine, but were still using Redis. Has that changed ?
Then I had a chance to re-build the system. I couldn't be more happy with it. It scales well, it's pretty fault proof and almost as fast as my EventMachine solution. Threads are slow, but the real slowness comes from the IO. Now instead of em-http-request I have much better working Curl::Easy and the whole codebase is much easier to understand.
"Sidekiq is compatible with Resque. It uses the exact same message format as Resque so it can integrate into an existing Resque processing farm. You can have Sidekiq and Resque run side-by-side at the same time and use the Resque client to enqueue messages in Redis to be processed by Sidekiq.
At the same time, Sidekiq uses multithreading so it is much more memory efficient than Resque (which forks a new process for every job). You'll find that you might need 50 200MB resque processes to peg your CPU whereas one 300MB Sidekiq process will peg the same CPU and perform the same amount of work. Please see my blog post on Resque's memory efficiency and how I was able to shrink a Carbon Five client's resque processing farm from 9 machines to 1 machine."
We recently released jobco, a simple Resque distribution. You can check it out at https://github.com/mrzor/jobco
We attempted to address some of the issues we had with our resque 1.x stack regarding productivity (ease to use both in development, and ease to deploy) and regarding API consistency (rewrote resque-status to keep the Resque API intact).
It doesn't address any of the high throughput issues that were mentioned earlier.
I'm definitely up to help fixing stuff within Resque when appropriate. However, the Jobco project is possibly more appropriate if the feature you're looking after can be provided using plugins, or involves providing better CLI interface to Resque.
You can find out my email address using `git log` on the jobco repo.
Numerous past articles covered this in one way or another.
Choosing between SideKiq and Resque is primarily a matter of choosing between forked processes and threads. Retiring Resque and let SideKiq be the present and the future would destroy that choice.
SideKiq, as a threaded impl., can claim some cool benefits : - shorter job launch times (only interesting if you actually run thousands of second spanning jobs a minute) - reduced memory footprint (you probably have little to no amount of thread local storage)
I will venture that the reduced memory footprint is mostly the result of the unfriendly to copy-on-write GC we have to live with using 1.9.
Resque has a clean and simple API and a decent ecosystem. It can and hopefully will evolve into something cooler and more powerful - even if worthy alternatives exist.
There are definitely a lot of people who are not happy with the way that Apple has been acting as of late. What if we were to select one (1) laptop model and work on making Ubuntu on it a seamless experience?
alrs is true ; in most cases your work is automatically replicated to Ubuntu (where the package-maintainer relationship is less strong than in Debian. Most packages don't have a dedicated maintainer there).
And everyone can do it ; no sign-up or complicated process required. You just have to find a sponsor for your packages which is just a fancy name for "every code needs review". Debian is often described as bureaucratic but most of the processes are very sane.
Canonical need to be more profitable, indeed they're probably not or barely profitable yet. They seem to have two ways of increasing profit a) monetising their existing desktop user base and b) Getting more desktop users so that more businesses become comfortable with it and buy business support.
Then what is the problem exactly? Every time Ubuntu comes up, people ask if it is fully hardware compatible. Is that just a branding issue, where people think Ubuntu isn't hardware compatible with most modern laptops?
There are multiple problems, the main one is that windows is known, comes "free" and already installed when you buy a computer, is less of a risk and is good enough. Most people will know a "computer guy" who can help them with windows or be comfortable that they can pay someone otherwise.
Software is a bigger problem in some ways than hardware now (though the web is fixing this), note the effort Ubuntu are putting into the software centre, allowing paid for apps, proposing to relax the restrictions to get new software into Ubuntu, running app competitions and so forth.
Selling operating systems isn't easy, look at how long apple have toiled in the wilderness and they have linked hardware, loads of good will via ipod/iphone, billions in the bank and have managed to climb to something like 5%-10% market share.
In contrast, Mozilla gets most of its income in an almost identical way (making google a default search provider) and that's not considered desperate.