Pulsar vs. Kafka
streamnative.io
streamnative.io
https://jack-vanlightly.com/blog/2018/10/2/understanding-how...
https://jack-vanlightly.com/blog/2018/10/21/how-to-not-lose-...
https://jack-vanlightly.com/blog/2018/10/25/testing-producer...
https://jack-vanlightly.com/blog/2019/9/4/a-look-at-multi-to...
They lost me here. I can think of plenty of situations where reduced latency is beneficial, but not many situations where shaving a few milliseconds would make a difference to “business insight”!
Although I suppose it is strictly correct, in the tautological sense...
Of course, our business is still running on nightly reports. But the tech is cool!
So I wanna say you're wrong, but I've got man-years invested in a system that hasn't been utilized to its full potential yet :(
Time will tell. Somewhere, some competitor will be using real-time insights to out-compete us.
Why? Can a business mobilise in anything less than days? If a report is minutes out of data, is that any loss? Given that some largish proportion of reports are never used, perhaps better management is key.
not disagreeing but efficiency is not just a matter of quickness.
This generally has implications on data quality, because you aren't fixing data quality issues once they occur in prod.
This also makes it far easier to develop ETL pipelines. You don't need complex tooling to see whether your ETL pipeline works.
You can technically fix data quality and dev velocity issues without low data freshness, but a quick glance at the data engineering landscape tells you they aren't being solved enough.
Quite simply, the end goal is increased oil production and reduced power consumption. Realtime data flows are a big deal when reading from thousands of sensors for just one process system and there are hundreds of tunable parameters. All this has to be safe also.
Bookkeeper has/had even worse documentation and setup. When I used it bookie couldn't have mounted directory for log storage, so it's not really that persistent as they say. Before restarting bookie node it has/had to be reformatted to allow bookie instance to re-use logs saved on disk. The whole "log are persisted" is true, but they don't say that you can't simple mount them in docker and restart your PC to it.
Pulsar is good when you get it working. Documentation is really bad and it's hard to make it work. All the super-positive articles about pulsar are sponsored (streamnative and yahoo) and biased to make pulsar look much, much, MUCH SIMPLER than it actually is.
Kafka clients can start at a particular offset ID within a partition.
It's not all hype either, according to this post https://jack-vanlightly.com/blog/2018/10/21/how-to-not-lose-... it seems like a solid piece of tech.
Then you combine that with how immature the UX is, and it really just doesn't feel good to deal with. It's not that much work to run a Kafka cluster with the sort of design that MSK provides, and it can be done better than that without much cost.
All the "success cases", sample architectures and real in the wild systems I've seen built on top of Kinesis have 1-3 consumers max.
They added the enhanced fan-out to try to get around this but it seems like you have to (over)pay for a ton of provisioned capacity to ensure you get decent latency on a > 5 consumer use case. So much like Lambda, it's only "serverless" or non-provisioned for people who have super low expectations of what system software ought to be capable of. For everyone else, it's just a mediocre expensive provisioned solution.
I see the same with Spark vs Flink in that similarities outweigh differences. I wonder if this is some sort of emergent pattern in open source software.
The only thing I can see that can make this true is Pulsar seems to have better elastic scalability. But it seems to score less on everything else. It has a much more complex storage system that ends up not matching Kafka's high-end throughput at large scale.
From what I recall, Twitter ended up abandoning BookKeeper due to storage scale concerns. Related: https://blog.twitter.com/engineering/en_us/topics/insights/2...
Pulsar likely would have been considered if it was more mature at the time and sported a community of comparable size to Kafka (it's still a long way from this).
1. A single partition is stored in one node (replicas on another nodes). With this, introducing new nodes takes very long time to replicate large partitions, because it can replicate one partition from only one node (leader of the partition). On Pulsar each segment of partition is stored in a different bookkeeper node.
2. Because of 1, if two consumers read different parts of a partition that are far from each other, they will compete over disk bandwidth. In Kafka consumer can not read from replica node. If a topic is really popular and many consumers try to read from it (from different parts of the file which makes OS page cache useless), total consumption rate is limited to disk bandwidth of a single node. But in Pulsar each consumer can read from different brokers. Catch up consumers won't trash streaming consumers in Pulsar.
These are not problems that can be fixed easily. Additionally, in the realm of streaming the difference between Flink and Spark is day and night. The low watermark feature that Flink offers makes them behave fundamentally different.
2. Kafka can read from a replica node. It's relatively new but it's there.
Topics (actually bundles of topics, called bundles) are what is assigned to Brokers. Topic assignment is dynamic, so when a new broker is added, the system will try and shed load from the busiest brokers to even it out on the system.
But unlike Kafka, when a topic is assigned to a broker, it doesn't have much state to move, mostly it just gets metadata added to it and opens a new "ledger" (which is just a chunk of the topics data over a time window, only one ledger is ever open at once). When it needs to serve data, it pulls that from bookkeeper nodes from previous ledgers, so the process of re-distributing load is pretty quick, it also doesn't eagerly pull in a cache.
Now, as far as the cache, that is primarily for "tailing reads", meaning, as writes occurs, and clients who are close to the tip of the recent data will just get it from the broker, without a need to pull it from bookkeeper. This is is one of the key parts about how Pulsar has multiple tiers of storage that help it have such good consistent latency.
Beyond processing writes, the biggest thing brokers do is handling "tailing reads" i.e., clients are consuming right near the tip of the topic. , this is the cache referred to. That means that when a new pbroker is three purposes:
1. Handling writes
Well, the Pulsar broker is (kinda) stateless, because they are essentially a caching layer in front of BookKeeper. But where's your data actually stored then? In BookKeeper bookies, which are stateful. Killing and replacing/restarting a Bookkeeper node requires the same redistribution of data as required in Kafka’s case. (Additionally, BookKeeper needs a separate data recovery daemon to be run and operated, https://bookkeeper.apache.org/archives/docs/r4.4.0/bookieRec...)
So the comparison of 'Pulsar broker' vs. 'Kafka broker' is very misleading because, despite identical names, the respective brokers provide very different functionality. It's an apples-to-oranges comparison, like if you'd compare memcached (Pulsar broker) vs. Postgres (Kafka broker).
Just to add to this, ease of use/setup is also a huge factor. There are technologies I can just spin up with zero knowledge and learn as I go. These are huge factors in adoption especially with Golang and nodejs.
Maybe it's a bit faster or a bit more elastic, or whatever, who knows. What I really care about is whether I get called at 3am and in that regard the argument seems pretty weak. Kafka for all its woes is a solid system you know you can count on.
I'd much rather see someone come up with a truly innovative alternative that actually pushes the boundaries, rather than just copying what's there already, and adding a few window dressings.
The page: https://pulsar.apache.org/powered-by/ suggests there's quite some number of corporate users who are happy to confirm they use this suite. I don't know how many of those you've heard of, though.
I suspect many private & government agencies around the world would decline to formally attach their name to any list like this, lest it be (mis)interpreted as an endorsement.
Looking forward to trying Kafka again when they finally remove ZK
That said, we have used Kafka for so long, and it works well enough at our scale that we have no reason to even test Pulsar. Kafka also has much more integrations with other tools due to its popularity.
Say you have a front-end dealing with the clients in a streaming manner (be it websockets or SSE). All front-end instances send messages to a topic on a messaging system. Processing is done with Flink or Spark, but now you need to get some answer back (or publish regular updates) to the client; so you push it to another topic on the messaging system. Works fine if you have a fixed and low number of front-ends; they pull everything and select messages for their clients. If you have more front-ends you want to have them pull only the messages destined to their clients. You might want to use Kafka partitions to do this, but it is kinda clumsy.
Furthermore if you need to scale the front-end, you'll have to reassign a partition scheme to all the front-end instances while they continue to cater to their specific clients. On top of restarting the Flink/Spark processing to fit the new partition scheme. I don't know of a simple way to do that with Kafka.
In Pulsar, the problem becomes _much simpler_: have the front-end chosse a UUID that represents them, send it as part of the messages, and interpret it as a return adress. The processing then pushes out to topics like: persistent://domain-x/app-y/back-to-clients-<uuid>. Done. No need for repartioning or topic creation.
Other than that, the pros are: the messaging Key_Shared mode [2], worth looking at; and you also get some message acknowledgement features. Cons is deployment, which is quite involved.
[1] https://pulsar.apache.org/docs/en/concepts-messaging/#no-nee...
[2] https://pulsar.apache.org/docs/en/concepts-messaging/#key_sh...
I’ve implemented exactly what you talk about using dynamic topics in Kafka and it was trivial.
Maybe I’m missing something?
Now with that said, these are the current downsides of Pulsar as recently perceived by me: More complex architecture (in the Kubernetes context) (zookeeper + broker + bookie + proxy + autorecovery [+bastion/prometheus/grafana]) vs Kafka's (zookeeper + broker [they are working on removing zookeper now]). Pulsar has extremely active development, but not that many active developers or active community size, large number of active bugs, the project is far less mature than Kafka, the documentation body is much smaller and worse, and the respective Stack-overflow (etc.) knowledge-base is much smaller than Kafka's. My experience with Pulsar's deployment to Kubernetes was that is that it wasn't ready. E.g. a lot of startup synchronization is being done by k8s yaml-embedded startup scripts and is extremely brittle and in some cases broken (you have to manually restart the proxy pod on startup, etc.). Due to the metric exposure on Kafka I find some critical devops scenarios (e.g. handling a full node loss) to be more transparent and somewhat more possible to approach than with Pulsar which appears more complex and more of a black box to me in this respect at this time. One big advantage often brought up in Pulsar vs Kafka comparison is that Pulsar has active sharding rebalance, but to my knowledge Kafka now has something like that too, albeit maybe not as dynamic as Pulsar has, not sure.
So to summarize, Pulsar is great, but actively developed and quite complex given the size of the (serious) active user base and the active developer base. I think the active user base being the main developmental driving force, it's a chicken-egg issue which simply takes its own time to evolve in parallel. We have to realize that Kafka is a 10 year old open source Apache product, thus very mature, and that's why I'd recommend it for new projects (which need to reach production quickly) over Pulsar at this time.
For a new project, you should try to document different aspects of requirements. Is it data streaming, queuing, or both? What's the data retention policy? Message rate? How many consumers and producers? Any inbound or outbound integration with 3rd party destination (i.e. S3, Flink)? Both Kafka and Pulsar have so many features to offer. It is not a simple task to pick one vs another. If you ask for guaranteed delivery, both will satisfy that requirement. A level up question would be who can guarantee in-order delivery.
Managing Kafka and Pulsar require knowledge. I do not think any of these durable messaging software is maintenance free (or industry is not there yet). Any reliable distributed system is complex out of necessity. These system more or less require log consensus algorithm to achieve high availability. They all use either zookeeper or one of raft implementations requiring multiple nodes to perform leader election. This is common in all distributed architecture (kafka, Pulsar, Cockroach, etcd...). I would attest Pulsar can be administratively simpler than Kafka, because of separation of broker and bookkeeper (data persistent layer). But this does not mean any dev-op without knowledge can proficiently manage the cluster. We use Kubernets/Helm to manage all of our Pulsar clusters. I would not credit Pulsar alone with low operation upkeep. It is combinations of Kubernetes, Helm, in-house tools, and engineering knowledge to lower the operation cost.
I use Kafka and needed a cache based on incoming events which were partitioned.
But if the consumer crashes it's not easy to pick up from exactly where it left off without manually committing offsets which hurts performance. There's also some hand waving Kafka gossip that it's hard to commit offsets right. Further suppose a task consumes events from topic A partition 3, produces events to topic B. Again, if the task crashes it's not clear what needs replaying in order to not to lose messages (on write) or missing messages (on read).
Any insight here is greatly appreciated.
See this Dec 2019 presentation by Pulsar committers, where they explain all this in more detail, i.e., the lack of transactions and the resulting limitations, and the motivation for adding such transactions to Pulsar. The approach looks very similar to Kafka's. https://www.slideshare.net/streamnative/transaction-preview-... The original ETA for transactions was Pulsar v2.6 (June 2020), but as of today there's still quite some work to be done (https://github.com/apache/pulsar/issues/2664). The latest ETA seems to be around the end of the year.
The key difference for an end user is that Kafka released all the functionality in one go back in 2017 (idempotent producer, transactions; which fwiw also explains why designing+building+testing took the Kafka community that long) so it has been much easier to understand what is actually supported vs. what is not.
For example, with Kafka Streams, any app you build with it just needs to set “processing.guarantee” to “exactly_once” in its configuration, and regardless of what happens to the app or its environment it will not lose messages (on write) or miss messages (on read) from Kafka.
Consider asking your question with a few more details in the Kafka user mailing list [1], or in the Confluent Community Slack [2] if you prefer chatting.
[1] https://kafka.apache.org/contact [2] https://launchpass.com/confluentcommunity
Give this a try: https://www.digitalocean.com/community/tutorials/how-to-inst...
However, the complexity you refer to is essential in a reliable messaging framework. Use of zookeeper or any log consensus algorithm requires multiple nodes ( 3 or more) to achieve durability and high availability goal. It is out of necessity. This is actually essential complexity.
There is another messaging framework called Nats.io. It is not persistent so architecturally relatively simpler. You might want to investigate.
Give me a break, I'd literally never even heard of Pulsar until this article popped up.
Of all messaging systems I would have thought Kafka vs SQS, or even RabbitMQ at the very least
Same - was it the comment saying nobody chooses Kafka anymore? I was surprised by it.
These people seem to be invested (quite literally) in Pulsar disrupting Kafka's position in the pub/sub & event streaming space. This isn't to say Pulsar can't someday be superior, though. Time will tell. But Kafka has such a massive lead and it fits the needs for most people who need to do data integration that I don't see the need for everyone to uproot their solutions in favor of Pulsar.
But almost everyone has heard of GNU/Linux or Apple OSX.
- a language for specifying data transformations, and
- an engine to compile programs written in that language into mapreduce jobs to execute on a hadoop cluster
it was designed to easily map some common functional and SQL idioms (e.g. filter, group by w/ aggregation functions) to parallel execution for processing huge amounts of data.
Impala is another big data project that is an engine for planning and executing SQL on data stored in a hadoop cluster.
Zookeeper is... black magic??
You don't need a CS degree to work in this field (I don't have one either!) but there are fundamental concepts you need to understand in order to make informed decisions when designing distributed systems.
Event driven concurrent framework for Python https://github.com/quantmind/pulsar
I still recommend managing Kafka through the CLI tools however.