Kinesis has a real auth story too, plus you can trigger Lambda functions off streams.
I read this quite often, but we run a relatively small kafka cluster on GCP and it's pretty much hassle-free. We also run some ad-hoc clusters in kubernetes from time to time which also works well.
What exactly have you found complex about running Kafka?
I run small 2-node kafka cluster that processes to 10 million messages/hr - not large at all - it's very stable, for almost a year now. However what was complex was:
* Setup. We wanted it managed it by mesos/marathon, and having to figure out BROKER_IDs took a couple hours of trial and error.
* Operations. Adding queues and checking on consumers isn't amazing to do from the command line.
* Monitoring. It took a while before I settled on a decent monitoring solution that could give insight into kafka's own consumer paradigm. Even still there are more metrics I would like to have about our cluster that I don't care to put the time in to retrieve.
Would you be willing to write a bit (or point to a post with) more about this? What do you find useful?
Secondly, due to the difference in speed in the consumer and producer, we typically have an offset lag of around 10MM, and its important to monitor this lag for us because if it gets too high, then it means we are falling behind (our consumers scale up and down through the day to mitigate this).
Next, we use Go, which is not an official language supported by the project but has a library written by Shopify called Sarama. Sarama's consumer support had been in beta mode in a while, and in the past had caused some issues were every partition of a topic wasn't being consumed.
Lastly, at the time we thought creating new topics would be a semi-regular event, and that we might have dozens of them (this didn't pan out), but having a simple overview of the health of all of our topics and consumers was thought to be good too.
We found Yahoo's Kafka Manager[1], which has ended up being really useful for us in standing up and managing the cluster without resorting to the command line. It's been great, but it wasn't exactly super obvious for me to find at the time.
Currently the only metrics I don't have are things plottable things like processed/incoming msg/sec (per topic), lag over time and disk usage. I'm sure these are now easily ingested into grafana, I just haven't had the time to do it.
All of this information is great to have, but requires some setup, tuning, and elbow grease that is probably batteries included in a managed service. At the same time however, this is something you get almost out of the box with RabbitMQ's management plugin.
In other words, I could probably get everything up and running (especially with the various Kafka-in-Docker projects I found), but what happens if (when) something goes wrong?
It's multi-tenant, but interaction is nearly identical to interacting with a dedicated kafka cluster -- i.e. you can use any regular kafka client library.
Check out docs[1] and launch blog post[2]. Happy to answer any questions here or through email (contact info in profile).
[1]: https://devcenter.heroku.com/articles/multi-tenant-kafka-on-...
That does not match my experience at all. Of all the distributed message queues I've tried, Kafka has been - by far - the easiest to operate.
It works well out-of-the box and even setting it up with ZooKeeper is relatively simple.
If that isn't you and you are still intent on using it, I'd second the opinion to go with "let someone else (e.g. 'cloud provider') administer it". It's probably not really worth the effort to get configuration settings and installation details right unless you genuinely have a serious interest in it.
We had a mix of realtime (continuous) consumers and batch consumers that consumed directly from Kafka, but also archived that data to S3 for future batch processing.
Operations were quite simple for us, but we were also used to running Zookeeper before, so YMMV. If you can get away with Kinesis, it's simpler operationally, but Kafka isn't difficult compared to most systems.
If it was a clean failure, no big deal. If the disk started having issues and taking a long time for IO to come back, all partitions with replicas or leaders on that node were slow. That's kind of the state of the world with most distributed things though: a bad node is worse than a dead node.
"no leader for partition" is something I see frequently.
In Kafka's defence a better managed cluster might do better. In nsq's defence: nobody has to manage it, it just runs.
Kafka is a distributed, partitioned log. Messages published to a partition in a given order will be stored in that order and replicated in that order, and consumed in that same order. Kafka writes all of your data to disk, and waits for replicas to acknowledge it (depending on your producer configuration). Kafka does considerably more work than nsq.
I like nsq, it works well, but they aren't at all designed to solve the same problems.
I mentioned the issues above because they'd happened, but in years of running Kafka it was never Kafka's fault.
> "no leader for partition" is something I see frequently.
Something is wrong with your cluster if it's constantly switching leadership. Leadership should be stable -- and furthermore, partitions have preferred leaders, so they should stay with specific nodes until that node goes down. Kafka has a process which shifts leadership back to preferred leaders automatically (unless it's off). Check your logs.
I can't remember if it's an issue for change capture since in theory you have some insert timestamp column anyway.
The best way to get the power, throughput, latency of Kafka without the operations is to use a hosted service. The one created and supported by the Kafka team is Confluent Cloud https://www.confluent.io/confluent-cloud/
Any time it seems like I have to establish some sort of relationship and get an enterprise tailored solution, I no longer feel like I'm part of a public and elastic market.
From what they said multi-tenant is about a month a way and self service cloud early next year.
Interesting to know about if I want to introduce Kafka in a bigger project, though.
Definitely not for a startup whose monthly Kinesis bill is a few hundred dollars, which is a shame because Kakfa is better in several ways.
[1] http://go.datapipe.com/whitepaper-kafka-vs-kinesis-download
I'd look at a hosted Kafka solution before going with Kinesis.
I'd still run it myself over Kinesis, even if I was a single-person startup.