Side note: https://pulsar.apache.org also seems to be gaining traction and has a much better story around performance, pub/sub, multi-tenancy, and cross-dc replication. Will be interesting to see the evolution of both going forward.
Side note: https://pulsar.apache.org also seems to be gaining traction and has a much better story around performance, pub/sub, multi-tenancy, and cross-dc replication. Will be interesting to see the evolution of both going forward.
Biggest difference is that Pulsar splits compute from storage. Pulsar brokers are stateless services that manage topics and handle clients (publishers and consumers) while reading/writing data from a lower tier of bookies, which is the Apache BookKeeper project for low-latency real-time storage workloads.
Better performance and easier scaling compared to Kafka's partition rebalancing. It also natively supports data center replication and multiple clusters working together globally. Pulsar addressing is property/cluster/namespace/topic so multi-tenant isolation is built-in.
Anywhere you use Kafka can use Pulsar and also consolidate pub/sub only systems that need lower latency than Kafka.
- rebalancing: Bookkeeper (which is the storage backend of Pulsar) manages rebalancing way better than kafka.
- Read scaling: Bookkeeper uses consensus to store log records, which means you can read from whatever shard you want. Whereas in kafka, reader has to connect to partition master.
The specific concern here is the possibility that Kafka's ISR strategy can potentially result in a corrupt leader partition and truncate messages to recover from a broker machine failure. The unclean leader election configuration setting for Kafka brokers is relevant here.
Also "better" is subjective depending on your configuration, requirements, and storage backends.
It's also been around for years but recently open-sourced so community is smaller than Kafka, however it's growing quickly along with the usual ecosystem of drivers, extensions, and services. Unless you're already running Kafka, I would look at Pulsar first for new projects.
A great overview of Pulsar history with plenty of links is here: https://streaml.io/blog/messaging-storage-or-both/
And pulsar does not provide exactly once neither, that is something very important.
- pulsar provides an failover subscription mode, which seems to be the equivalent of partition rebalancing of consumer group in kafka. https://pulsar.incubator.apache.org/docs/latest/getting-star...
- it has partitioned topics as well.
- it supports idempotent producing and have effectively-once delivery semantic.
It seems to have all the kind of primitives for kafka streams. to use.
Pulsar doesn't have any client overhead since it's all tracked on the broker so there's no real need for a separate library. Read, process, and push messages using any code you want, 1 at a time or in batches.
There's no realistic "exactly once" either, it's idempotency or some local cache of processed events used to dedup, you can read more here: https://streaml.io/blog/exactly-once/
“Kafka has exactly-once delivery but messaging system x/y/z doesn’t provide it” is also confusing and misleading. Exactly-once is technically effectively-once: “at-least-once” and make the processing of the messages idempotent or “de-duplicated”. This has already been done in the industry for many decades in many mature stream processing engines like heron, flink. It isn’t a really new thing. And many messaging systems like pulsar already provides those primitives (e.g. idempotent producing, at-least-once delivery) for processing jobs to achieve effectively-once very easily. Streamlio folks did a great job about explaining exactly-once and effectively-once. It is worth checking this blog post out -- https://streaml.io/blog/exactly-once/
I think Pulsar itself as a distributed messaging system does provides all the three delivery semantics: at-most-once, at-least-once and effectively-once. It is very easy for people to use and integrate. I don’t think it is difficult to make kafka streams run with pulsar technically. The question is more is there a value to do that, do kafka folks wanna to do that, can the collaboration happen in the ASF?
that's just my two cents.