I will take the opportunity to say that Kafka is kind of painful, with or without ZK. Check out NATS! [0]. It doesn't solve all the same problems, but is so much easier to use (during development especially) and can do a lot of the same things.
I will take the opportunity to say that Kafka is kind of painful, with or without ZK. Check out NATS! [0]. It doesn't solve all the same problems, but is so much easier to use (during development especially) and can do a lot of the same things.
NATS doesn't ever store messages persistently; but this might be fine for your application, and then you don't have to worry about setting 5 different config options to make sure Kafka actually frees up disk space like you expect it to ;)
NATS also enables some unique patterns like request/reply via a "reply to" message header.
Anyway, it's been a joy to use!
Not true. Both Nats streaming and the upcoming jetstream (core nats) do.
To the folks at Synadia -- I love NATS, but the naming and organization of these projects could use some work. What's with the `stan.*` repository names? Where did "jetstream" come from? Why is it baked into `nats-server` but `nats-streaming-server` isn't? Is `nats-streaming-server` on the back burner?
[0] https://docs.nats.io/nats-streaming-concepts/intro [1] https://github.com/nats-io/jetstream
Jetstream is GA with the 2.2.0 release. Folks who believe in waiting for "not .0" won't have to wait too much longer.
If you went there for tap water, yeah, maybe there are better options.
That being said, have you checked out NATS Streaming Server? It’s effectively a first party client for NATS that gives it at least once semantics and persistence, and makes it much more applicable to use cases that are currently on Kafka.
Docs here if you’re curious - https://docs.nats.io/nats-streaming-concepts/intro
[0]: https://docs.nats.io/whats_new_22 [1]: https://docs.nats.io/compare-nats
Disclosure: I work for Confluent
> messages from a given single publisher will be delivered to all eligible subscribers in the order in which they were originally published. There are no guarantees of message delivery order amongst multiple publishers.
https://docs.nats.io/faq#does-nats-offer-any-guarantee-of-me...
Messages are ordered within partitions.
> What systems out there require strictly ordered data? It seems like any design that requires something like that is going to be extremely brittle.
TCP/IP ?
Right, but that means you're still "unordered" across those partitions?
> TCP/IP ?
But TCP/IP isn't delivered in order, it rearranges the unordered packages by their ID. I guess ordered delivery would be nice for that, but I just feel like making your protocol not require ordering is far simpler.
Not to mention that both TCP and Kafka have to handle head of line blocking?
I'm not trying to say that ordering is bad or anything, I just feel like it isn't buying me tons.
Right, so related messages have an ordering guarantee but unrelated messages may be processed out of order relative to each other, which is usually what you want. (Of course you do have to set the record key correctly).
> I'm not trying to say that ordering is bad or anything, I just feel like it isn't buying me tons.
It's a lot more lightweight than full ACID, but if you get your dataflow right it achieves everything that a traditional database does. Without ordering you wouldn't be able to do anything that requires any kind of consistency.
To me, it seemed at odds with the parallelism of a partition, but I suppose in this case you'd be partitioning on some sort of semantic key vs, say, a hash.
Thanks for bearing with me on that, this was just an unfamiliar idea for me.
> both TCP and Kafka have to handle head of line blocking
Well which is it?
(If TCP doesn't give you ordered delivery, why would a head block the rest of the line?)
Maybe if you're an e-business, you'll split everything happening on your website by client id, but still want events belonging to a single client to be received in order, for practicality.
Apache Pulsar offers the same distributed log offering with a fundamentally better architecture, but Kafka has closed most of the gaps now and has far more integrations and a bigger ecosystem.
That being said, I don't think this is what differentiates the two systems, the guarantees they do/don't make are likely what will make the decision for your project.
[0]: https://github.com/nats-io/nats-streaming-server/blob/master...