It's worth bearing in mind that messages per second is important, but it's easy to get fixated on benchmark porn.
Different queueing systems have different guarantees, disciplines and semantics. These affect user-facing behaviour, which tends to be more important at first blush than throughput.
I like zillions of messages per second as much as the next fellow. But frequently you need to worry about things like:
* Are messages delivered once and only once?
* Are messages are messages delivered at least once?
* Can messages be dropped entirely under pressure? (aka best effort)
* If they drop under pressure, how is pressure measured? Is there a back-pressure mechanism?
* Can consumers and producers look into the queue, or is it totally opaque?
* Is queueing unordered, FIFO or prioritised?
* Is there a broker or no broker?
* Does the broker own independent, named queues (topics, routes etc) or do producers and consumers need to coordinate their connections?
* Is queueing durable or ephemeral?
* Is durability achieved by writing every message to disk first, or by replicating messages across servers?
* Is queueing partially/totally consistent across a group of servers or divided up for maximal throughput?
* Is message posting transactional?
* Is message receiving transactional?
* Do consumers block on receive or can they check for new messages?
* Do producers block on send or can they check for queue fullness?
And there's probably a bunch more I've forgotten.
The thing is that answers to these questions will fundamentally change both the functional and non-functional nature of your queueing system.
For example, a queue system giving best-effort, unordered, non-durable behaviour is going to run a lot faster. It also pushes a lot of work onto the application programmer. On the other hand, once-and-only-once, durable, consistent queues are lot slower and screech to a halt under most partition conditions. But they also fit what most application developers expect to happen upon the first encounter with queueing systems.
I work on a section of Cloud Foundry in my day job, and other teams have seen that different tasks require different queueing approaches.
For example, stuff like metrics is is still useful under conditions of dropping messages, out-of-order messages and so on, because what's interesting is the statistics, not any one single measurement.
But a message like "start this app" requires much higher guarantees of ordering, durability, delivery certainty. People get mad if your PaaS doesn't actually run the application you asked it to run.
So, just remember: queues are not queues. You need to compare delivered apples with lossy oranges.
As a note, the author observes that MQTT provides an option to select which delivery semantics you prefer (at-least-once, at-most-once / best-effort, once-and-only-once), but I can't see which one the benchmark is run for.