They specifically want Kafka, there's no real reason other than they need a queue, which Kafka actively states that it's not. At that point it gets really tricky to reason with the developers about why they might be better served by something else. Generally speaking it's not much of an issue, because Kafka will deal with workloads just fine, it's just weird. I have seen one customer use Kafka as a database, that works less well.
We do see the same with Kubernetes. The developers pick Kubernetes and at that point it's to late. They specifically want Kubernetes even if you could more easily solve the problem with Nomad, Docker-Compose, plain old VMs or EC2, depending on the problem.
I wonder if something similar is happening with Kafka?
As an engineering leader sometimes that even means knowing that people are making the wrong decision, and letting them do it anyway, and then helping them learn from it.
Confluent itself says that for a workload of 38Mbps/sec rabbitmq has 5 times less latency than Kafka and is 100x simpler to manage.
https://www.confluent.io/blog/kafka-fastest-messaging-system...
How do you teach someone to look at the problems first and then pick the tools?
By understanding what the person cares about. Everyone knows "pick the right tool for the problem". Not everyone uses such a simple calculus because life isn't that simple. People have their own agendas, backgrounds, experiences, career growth desires, personal lives, etc., that are all part of their personal objective function. If you want to convince someone that your tools are better, show that your tools have a higher payoff for their personal objective function. This is way more than a mere product question. In a team setting it's even harder, because you have to balance it across multiple people simultaneously.
"Due to CPU bottlenecks, we were not able to drive a throughput higher than 38K messages/s, and any attempt to measure latency at this rate showed significant degradation in performance clocking a p99 latency of almost two seconds."
Where do the docs state that? Might be a useful link to keep handy.
https://www.confluent.io/blog/kafka-fastest-messaging-system...
They do go to great length to avoid calling Kafka a queue. No where does it directly state that Kafka is not a queue. The docs just never talks about Kafka as being a queue.
The specific thing I have experience with is in analytics/relational databases. Suddenly around 8-10 years ago it became imperative for every client I was dealing with to migrate their RDBMSes to Hadoop/Hive setups, even when their largest dataset was only about 120M rows denormalized.
They were trading three servers (primary, backup, DR) for sometimes 15 to 20. Queries that MSSQL was handling sub-second were suddenly taking 45s on Hive. It was utter madness and was as far as I can tell driven by good salespeople, FOMO, and the feeling of importance of being able to say your company is running Big Data(TM?).
I saw maybe one implementation (of dozens) that actually stayed in use for any length of time.
If you're using it, you're constrained by its engineering decisions, so you need to be sure it's a worthwhile tradeoff.
When scaling out on a standard message queue you generally have greedy consumers which means you can't assume stickyness, or have to create your own partitioning structures. It makes it great for realtime apps...
If there were cheaper alternatives I think people would use them but it does definitely have powerful benefits.
We actually have a use-case that exactly matches this. One service makes [stuff], the other service consumes [stuff], both services are "immutable infrastructure" with no local stable state storage, and [stuff] is individually too small and frequent to be affordable with IaaS managed-MQ per-message costs — but batching messages into reasonable chunks before send means potentially losing up-to-a-batch worth of messages if the producer dies.