A lightweight, high-performance, language-independent job queue system
github.com
github.com
I appreciate any entry into the space but this is a really crowded space. With some very battle tested entries in the field as people have pointed out.
There are a few questions not in the readme that I think need answering for a queue system:
- Execution guarantees (at least once, at most once, exactly once?)
- Order guarantees (FIFO, approximately FIFO, nondeturministic, etc)
- Throughput compared to other systems
- Fault tolerance characteristics, how many nodes can I lose before it stops working, when it does how do I recover
As I said, I have a lot of choices in this field. I'd like to see all the data up front.
My biggest concern is its reliance on mysql. There is no way this could be a valid option for high volume messaging when it is essentially database as a queue.
> Beanstalk is a simple, fast work queue.
> Its interface is generic, but was originally designed for reducing the latency of page views in high-volume web applications by running time-consuming tasks asynchronously.
The protocol is easy to drive directly and there are good libraries for most common languages.
Beyond being a great piece of software, I find the protocol to be really well-designed for a work queue.
As other comments have pointed out, HTTP is not a great idea because timeouts, etc.
The one subtle benefit that I can relay is that by using your main database as the storage layer, you can enqueue tasks within a transaction, as per: https://github.com/thruflo/ntorque/blob/master/src/ntorque/c...
RabbitMQ also supports transactions for some queue operations, but its notion of "transaction" and what you can do inside of one is much more limited than a typical database's: https://www.rabbitmq.com/semantics.html
Is that because mysql doesn't have FOR UPDATE SKIP LOCKED?
EDIT: TIL Redis has the option to turn on fsync-to-disk on every write. Probably not what people are thinking of when they suggest Redis as lightweight.
Currently Alpha: https://cloud.google.com/sdk/gcloud/reference/alpha/tasks/
(I work for GCP)
https://cloud.google.com/appengine/docs/standard/java/search...
GCP and Elastic did partner to offer hosted Elasticsearch though it's not a fully managed service: https://www.elastic.co/about/partners/google-cloud-platform
Basically, if you have messages which can take different amounts of time to process, or you need to quickly dynamically scale the number of consumers on a queue in response to volume.
How well does the "update max_workers" queue-modification command work in situations of very high message volume and/or high consumer counts?
But, RDMBS only makes sense if you can use your existing installation, so MySQL-only is a nonstarter.
can it break when MySQL's thread ID is reused, and some new client will get the same CONNECTION_ID as a previous failed worker?
In answer to your question generally: push-based models, while more complex, tend to be higher performance (by dint of improving throughput: the broker can push messages to a consumer while it's working on other things rather than waiting for a "gimmie" request; the broker can also coordinate when and how it delivers messages to which consumers for maximum performance, which can lead to significant speedups in high throughput situations).
A very powerful pattern that gives a reasonable amount of control in a push-based situation is combining a push based model with client acknowledgements, and a client "window" of a number of messages that the client may or may not have noticed yet, especially if messages can be taken back from that window programmatically in the event of slow or dead consumers. This is what RabbitMQ/AMQP calls "Qos".
In my experience, push-based messaging models should typically not be adopted up front, unless throughput requirements are known to require such a model. The added complexity (mental and in code) of managing a push-based queue model is very rarely worth investing in at the beginning of development.
Furthermore, it's possible to have extremely high performance in a simpler pull-based model, provided you make some tradeoffs. This is what Kafka does.
I would recommend switching from pull to push only when it becomes necessary (though this can be a non-trivial amount of effort depending on how tightly integrated your code is with your messaging system). RabbitMQ/ActiveMQ/Redis/Resque as brokers will shine here, since they all support both pull models ("get" in AMQP) and push models.
I'm just struggling to imagine a use case where a push queue would be preferable over a pull queue. I'm sure they exist, I've just never encountered one before. Seems like the major difference is centralized throughput control, which would allow you to minimize variance in message processing latency. There are similar use cases in e.g. operating systems for minimizing latency variance for better UI responsiveness, but I can't think of any concrete use cases for using this in high scale backend queues.
[update] Oh, it's "job queue system", not a generic "queue system". So it's not quite a perfect fit, I guess :)
I'm not saying this project is bad - I have no way of knowing. It might have legitimate usecases where it excels. But what it advertises can't be true.
"Lightweight" is really only interesting to me in two cases: First, you're designing it for limited resources, like an embedded system, for which the standard answer is simple too large to even consider. Second, when the standard answer in the field is so "heavy" (an ill-defined term itself, but moving on) that it causes problems of its own. JVM solutions sometimes get to this point, where the act of administering the solution itself gets bogged down in merely administering the JVM.
I do not personally have the problem that my job queues are too heavy, nor have I heard anyone else complain that ZMQ or Redis are just so heavy for what they do.
Can I hear more about the rationale behind this?
Could it be used as a replacement to Kafka message queue/ring?
(me is still looking for a Kafka-like piece coded in Go)
Otherwise, the cloud native landscape has some similar projects listed in the same category as kafka: https://github.com/cncf/landscape