For a new project, you should try to document different aspects of requirements. Is it data streaming, queuing, or both? What's the data retention policy? Message rate? How many consumers and producers? Any inbound or outbound integration with 3rd party destination (i.e. S3, Flink)? Both Kafka and Pulsar have so many features to offer. It is not a simple task to pick one vs another. If you ask for guaranteed delivery, both will satisfy that requirement. A level up question would be who can guarantee in-order delivery.
Managing Kafka and Pulsar require knowledge. I do not think any of these durable messaging software is maintenance free (or industry is not there yet). Any reliable distributed system is complex out of necessity. These system more or less require log consensus algorithm to achieve high availability. They all use either zookeeper or one of raft implementations requiring multiple nodes to perform leader election. This is common in all distributed architecture (kafka, Pulsar, Cockroach, etcd...). I would attest Pulsar can be administratively simpler than Kafka, because of separation of broker and bookkeeper (data persistent layer). But this does not mean any dev-op without knowledge can proficiently manage the cluster. We use Kubernets/Helm to manage all of our Pulsar clusters. I would not credit Pulsar alone with low operation upkeep. It is combinations of Kubernetes, Helm, in-house tools, and engineering knowledge to lower the operation cost.
Now with that said, these are the current downsides of Pulsar as recently perceived by me: More complex architecture (in the Kubernetes context) (zookeeper + broker + bookie + proxy + autorecovery [+bastion/prometheus/grafana]) vs Kafka's (zookeeper + broker [they are working on removing zookeper now]). Pulsar has extremely active development, but not that many active developers or active community size, large number of active bugs, the project is far less mature than Kafka, the documentation body is much smaller and worse, and the respective Stack-overflow (etc.) knowledge-base is much smaller than Kafka's. My experience with Pulsar's deployment to Kubernetes was that is that it wasn't ready. E.g. a lot of startup synchronization is being done by k8s yaml-embedded startup scripts and is extremely brittle and in some cases broken (you have to manually restart the proxy pod on startup, etc.). Due to the metric exposure on Kafka I find some critical devops scenarios (e.g. handling a full node loss) to be more transparent and somewhat more possible to approach than with Pulsar which appears more complex and more of a black box to me in this respect at this time. One big advantage often brought up in Pulsar vs Kafka comparison is that Pulsar has active sharding rebalance, but to my knowledge Kafka now has something like that too, albeit maybe not as dynamic as Pulsar has, not sure.
So to summarize, Pulsar is great, but actively developed and quite complex given the size of the (serious) active user base and the active developer base. I think the active user base being the main developmental driving force, it's a chicken-egg issue which simply takes its own time to evolve in parallel. We have to realize that Kafka is a 10 year old open source Apache product, thus very mature, and that's why I'd recommend it for new projects (which need to reach production quickly) over Pulsar at this time.
I use Kafka and needed a cache based on incoming events which were partitioned.
But if the consumer crashes it's not easy to pick up from exactly where it left off without manually committing offsets which hurts performance. There's also some hand waving Kafka gossip that it's hard to commit offsets right. Further suppose a task consumes events from topic A partition 3, produces events to topic B. Again, if the task crashes it's not clear what needs replaying in order to not to lose messages (on write) or missing messages (on read).
Any insight here is greatly appreciated.
See this Dec 2019 presentation by Pulsar committers, where they explain all this in more detail, i.e., the lack of transactions and the resulting limitations, and the motivation for adding such transactions to Pulsar. The approach looks very similar to Kafka's. https://www.slideshare.net/streamnative/transaction-preview-... The original ETA for transactions was Pulsar v2.6 (June 2020), but as of today there's still quite some work to be done (https://github.com/apache/pulsar/issues/2664). The latest ETA seems to be around the end of the year.
The key difference for an end user is that Kafka released all the functionality in one go back in 2017 (idempotent producer, transactions; which fwiw also explains why designing+building+testing took the Kafka community that long) so it has been much easier to understand what is actually supported vs. what is not.
For example, with Kafka Streams, any app you build with it just needs to set “processing.guarantee” to “exactly_once” in its configuration, and regardless of what happens to the app or its environment it will not lose messages (on write) or miss messages (on read) from Kafka.
Consider asking your question with a few more details in the Kafka user mailing list [1], or in the Confluent Community Slack [2] if you prefer chatting.
[1] https://kafka.apache.org/contact [2] https://launchpass.com/confluentcommunity
Say you have a front-end dealing with the clients in a streaming manner (be it websockets or SSE). All front-end instances send messages to a topic on a messaging system. Processing is done with Flink or Spark, but now you need to get some answer back (or publish regular updates) to the client; so you push it to another topic on the messaging system. Works fine if you have a fixed and low number of front-ends; they pull everything and select messages for their clients. If you have more front-ends you want to have them pull only the messages destined to their clients. You might want to use Kafka partitions to do this, but it is kinda clumsy.
Furthermore if you need to scale the front-end, you'll have to reassign a partition scheme to all the front-end instances while they continue to cater to their specific clients. On top of restarting the Flink/Spark processing to fit the new partition scheme. I don't know of a simple way to do that with Kafka.
In Pulsar, the problem becomes _much simpler_: have the front-end chosse a UUID that represents them, send it as part of the messages, and interpret it as a return adress. The processing then pushes out to topics like: persistent://domain-x/app-y/back-to-clients-<uuid>. Done. No need for repartioning or topic creation.
Other than that, the pros are: the messaging Key_Shared mode [2], worth looking at; and you also get some message acknowledgement features. Cons is deployment, which is quite involved.
[1] https://pulsar.apache.org/docs/en/concepts-messaging/#no-nee...
[2] https://pulsar.apache.org/docs/en/concepts-messaging/#key_sh...
I’ve implemented exactly what you talk about using dynamic topics in Kafka and it was trivial.
Maybe I’m missing something?