Show HN: goworker – a faster, Resque-compatible background worker
goworker.org
goworker.org
I would appreciate all kinds of feedback, and for anyone who wants to try goworker out, if you give me your email at the bottom of the page, I'll send you my contact info for some free one-on-one help getting started.
I should really document it. It has full type safety, stats, scheduled jobs, error reporting, retries, etc.
Full Sidekiq compatibility (scheduling, retries) was my next goal.
Please let me know how I can help you stay abreast of what we're doing with Resque 2.
Wondering how these ruby/rails inspired work queues compare with NSQ.
Author of go-workers here. On SIGTERM, I stop accepting new work, and wait for all running workers to finish before halting.
If workers take longer than the 10 seconds Heroku gives you, go-workers uses reliable queueing (using http://redis.io/commands/brpoplpush) so the job will run again next time you start up the process.
We're keeping our outbound work in sidekiq for now (less throughput; not as necessary to have fast deserialization). It would be interesting to move all the workers to Go (and explore the advantages of the libraries linked here) as well.
Have you also tested the performance of this with 40-50 workers doing intensive I/O work such as crawling web pages? One of the main disadvantages of Sidekiq is that the workers just freeze every 15-30 minutes when you have a ton of workers crawling web pages. The only workaround for me is to setup a cron job that restarts the workers if I detect this pattern.
On the queuing side, you use Ruby. For the examples, I was trying to show what the "equivalent" Ruby worker was so you could get an idea of what kind of workload it was. However, I do see that it is confusing.
I suppose the main use case is working on data loosely coupled with web app, thus ? For example, if we have to do any big processing on active_record objects, we would have to reimplement all model stack in go ?
I'd much rather see neat algorithms and libraries being re-written in C with an FFI-friendly API. I think it would reduce the number of "Show HN: I rewrote X in Y" posts. The majority of HN users rarely care about Y. Such posts are just attention-grabbing noise.
If it was written in C you could just write bindings to it in Y and nobody but Y developers would need to hear about it. Instead we could read about X on HN. X is interesting.
It doesn't seem right that we can't leverage libraries in other languages without significant effort.
I've been using jesque to consume Resque tasks with no problems, and performance is great.
regarding goworker, so the worker logic has to be written in go correct? for my rails app if its something simple like sending email, I can use that, but if it involves something where I need activerecord is it still a good option?
In real world use cases, how often is that kind of overhead a significant portion of total run time, compared with the actual work? If it's a small portion, would that make the speed up fairly irrelevant, does it actually speed up real world use cases significantly?
Of course, your actual work payload may be faster too in go than ruby if you know how to write it well.
I wish more people paid attention to questions like this. Be it web servers, app servers, whatever; too many people look at microbenchmarks and decide that a move is imperative when they're really only optimizing a tiny fraction of the total run time.
> Of course, your actual work payload may be faster too in go than ruby if you know how to write it well.
I think that this is the big bet you're going to make when you choose goworker over a Ruby implementation.
What makes this intriguing is that in a web app scenario, background workers are often used for tasks that would take too long to run in the HTTP request/response cycle. This means there is a selection bias toward tasks that may be computationally difficult. That's not to say that all tasks handed off to background workers are CPU bound; many (most, maybe?) are I/O bound, but for the class of problems that are CPU bound, this offers a great solution.
The tool-selection argument quickly devolves in to the usual arguments between the benefits of languages like Ruby versus lower level languages, but I'm not sure that's the most productive conversation. That's well traveled ground.
What I'm saying is that I'm happy to see this (Golang workers available through a Resque API) as an option. I'm not advocating blindly replacing C implementations (or any other implementation), but I wouldn't be surprised to see a Golang back-end outperform Ruby code, even when that Ruby code is calling in to C libraries. The usual rules apply when making that evaluation. You have to see for yourself.
It seems like using something designed from the start for cross-technology communication (like 0mq) may be a better idea, but I suppose there is already lots and lots of existing resque client code out there.
You can replace one single Resque worker to see how the performance improves without changing the front-end queuing or having to run an additional queuing server
However, even when not comparing (seconds to computer 10000 jobs) benchmarks, you still get benefits from goworker. Since dequeuing, deserialization, and updating stats is faster, the latency between "job insert" and "job started" is lower, which means faster interface updates.
You can also restructure your workers. Instead of running one worker that loops 5,000 times because the Redis overhead is too high, you can have one job queue 5,000 more which can be run on several distributed workers.
(I am confused with your first paragraph and a half, I think you have some typos in there, or I just don't understand what you're saying. But I understand what you're saying about latency and about being able to restructure tasks to be smaller, thanks.)
I mean that if your users are waiting on jobs, they get the result when `encode time` + `enqueue time` + `wait time` + `dequeue time` + `decode time` + `execute time` is completed.
Even though you might not need raw 1000 jobs/second throughput, your users will still benefit by reducing `wait time` + `dequeue time` + `decode time`, which goworker does.
Say a user clicks a link that creates an async job that does a API call somewhere. What's a good way to test that entire process at once? You can easily test them separately, that clicking a link creates an async job, and that the worker can process the job, but it's useful to test the whole system at once.
Same question goes for any time of queue system, like quque_classic, which stores the jobs in a postgresql table.