> This problem was never root caused, though we suspect an issue in the Celery workers themselves and not RabbitMQ.
So the problem wasn't Rabbit itself, but the usage/lack of understanding.
It's a great writeup. But I'm personally curious about them hitting the limits of vertical scalability. They said they used the biggest node 'available to them' without clarifying how big that was.
If the issue was RabbitMQ hitting resource limits then upgrading its hardware (or swapping it out for some other similarly featureful broker but which is faster) might have fixed some of the problems too, albeit not things like insufficient observability.
I also wonder to what extent RabbitMQ suffers from being written in Erlang. Switching to a featureful broker with sharding and replication that isn't Kafka e.g. Artemis might have allowed them to avoid rewriting code to not use helpful features.
This performance blog suggests RabbitMQ may top out at ~4000 messages/sec when making things durable, which isn't especially good performance:
Would I personally reach for kafka out of the box? Not so much. I personally think it is quite a monster to get setup correctly. If you need dual datacenter up time it also becomes very expensive as you need to buy cloudera or confluent. Or have a plan to manage it yourself which will be 'interesting'.
MirrorMaker is pretty decent and part of the oss release. There are also multiple alternatives available. There is no need to buy anything.