KIP-932: Queues for Kafka
cwiki.apache.org
cwiki.apache.org
On the other hand, Kafka isn't the only player in the queue game nowadays. If you need message queue and job queue semantics combined (which you likely do), just use Pulsar.
I'm just hoping librdkafka gets good too-tier support for this feature in a timely manner.
RabbitMQ has implemented Streams and "Super Streams":
> Super streams are a way to scale out by partitioning a large stream into smaller streams. They integrate with single active consumer to preserve message order within a partition. Super streams are available starting with RabbitMQ 3.11.
I wouldn't use Kafka for a job queue, and wouldn't use RabbitMQ for streaming data when ordering would be important.
Kafka is a stream, RabbitMQ is a queue. Without getting into the details, RabbitMQ is designed to add things to a stack and pop them off when consumed. Kafka is designed to stream everything to a continuous log and anywhere can tune in when appropriate.
With Kafka it "just" keeps appending to a dumb (but huge) circular buffer. But you can have multiple consumers read off this buffer and can start any point. Downside is customers have to maintain their own offsets (in some storage) but there is now a big decoupling between producer and consumer. This contributes a large part to high throughput too (and consumers can go at their own pace -ofcourse if they are too slow they can fall off the log).
Minor correction: you can maintain the offsets yourself if you want, but usually it's not necessary because Kafka can do it for you.
The abstraction Kafka provides is that for each consumer group and for each (topic, partition) tuple, your consumer object that is guaranteed not to receive messages before the last offset at which you called commit(). Internally, the committed offsets are stored in a special Kafka topic of their own.
https://eranstiller.com/rabbitmq-vs-kafka-an-architects-dile...
RabbitMQ messages are supposed to be processed/consumed/acked only once. Your app most probably won't ever get two exactly the same messages, unless you misconfigured/misued RabbitMQ. It's good for classic message processing - "used clicked something, run the job of informing subscribers that new post has been created" (because you can't send 1000 messages from a web worker thread).
Native capacity to do queueing is exciting, especially the concept of share groups to allow potentially different types of queues and shares in the future.
It’s might not be appropriate, but one step closer to eating more workflow engines for lunch.
Deno Queues - https://news.ycombinator.com/item?id=37674752 - Sept 2023 (67 comments)
RabbitMQ vs. Kafka – An Architect’s Dilemma (Part 1) - https://news.ycombinator.com/item?id=37574552 - Sep 2023
Is is the ability to parallelize consuming to a super large amount, without having to setup partitions ?
I'm planning on using kafka for a job queue, and i like the idea that i can add ordering definition if i need to, and that i can keep the jobs in the queue for auditing later on if i need to. What am i missing ?
Without being limited by partitions. In Kafka your unit of parallelism is partitions but what happens when you don't care at all (or much) about ordering and just want to add or remove consumers to match your current load? Queue semantics.
In Kafka the number of partitions can go up, but not down. And even when you do that the messages don't get split up to fill the new partition so you can't burn down a backup by adding more partitions or more consumers -- ope.
Kafka is an amazing event streaming platform.