~2M msgs/second messaging system written in Go
gist.github.com
gist.github.com
If it's something like 1 Gig, then it's the OS that should be commended for the miracle of throughput, not Go. Even VB would be able to pull numbers like these with aggressive buffering.
A more sensible metric would be to measure the throughput and the longest time in transit. If you get Go deliver 2 mil/sec with sub-ms delivery time, then we'll have something to talk about.
$ time seq 5000000 >/dev/null
1.10s user 0.00s system 99% cpu 1.107 total
There, my Mac Mini can push ~5M msgs/second!!1oneDo I get a pony now?
Seriously, what is this doing on HN and what on earth are people discussing?
If you want to brag with benchmarks then how about providing at least a remote clue about what you are measuring...
I'm looking forward to the full suite being released so we can repeat the tests and see how it fares under real-world conditions with different message-sizes etc.
Sorry for the snark in my initial comment but I found this gist really lacking (akin to those press releases where $vendor brags about some arbitrary figure without providing any details).
I do like the simplicity of NATS and look forward to a fast server implementing it. However for something like a MQ where performance is the key-metric and highly dependent on the chosen workload one really shouldn't throw around numbers without backing them up thoroughly - that only hurts credibility.
Please take the write-ups by the RabbitMQ guys as a guide, who publish the source-code for all their benchmarks and go to great lengths explaining them:
http://www.rabbitmq.com/blog/2012/04/25/rabbitmq-performance...
No, but you get an award for snarky 4-chan like behavior.
I wrote a system that tried to push as many messages through a socket as possible, whether it was tcp, udp, unix domain, or posix message queue. The goal was to determine which IPC mechanism was best suited for highest concurrency rather than highest throughput, but I think the results are interesting for both.
I wanted to measure the amount of time it took to do the same amount of work (i.e. enqueue and dequeue one million items) for each IPC mechanism and how concurrency affected the performance. The system uses a single producer with a varying number of consumer processes.
The results are available for browsing here: http://queueable.herokuapp.com/
The stream socket implementations (tcp, unix domain stream socket) actually perform a minimum of two write()'s per queue item - once for the length and once for the actual content. Both of these are wrapped in loops to ensure the full content is written, so occasionally more than two write()'s might occur.
On one of the linux hosts, the TCP implementation can push 1 million messages through in roughly 0.66 seconds using a single producer thread, which corresponds pretty closely to the 2 million messages/second claim for NATS. The POSIX message queue can do it 0.48 seconds, which corresponds to more than 2 million messages/second, but POSIX messages queues are a datagram implementation that only require one mq_send() per message.
I think this shows that 2M syscalls/second is indeed possible. I made no special effort to optimize the C++ code. If you'd like to review the code or run the tests yourself, feel free to check out the code at https://github.com/adamonduty/queueable . Anyone can run the tests and submit results to view on the web interface.
Oh, and interestingly, the Macbook Pro I used to generate OS X results was the most recent hardware but among the slowest in wall-clock performance. OS X also shows a zig-zig effect as message size increases, similar to SunOS. Not sure why.
Note that the experiment was done on an MBA.
Of course, latency will suffer. You can't have both max throughput and min latency. It's a choice to be made.
[1]: assuming 4Kb packets
Ultimately, if you want real speed doing something real simple, you're going to want to use FPGAs anyway which would be a trivial consulting fee to implement compared to a development effort that will go deep into your kernel, hardware drivers and probably protocol tuning as well.
I think the main utility for such a benchmark though is to establish a lower limit on theoretical per-message overhead. Any practical system is likely to want to do something interesting with the content of the messages.
But this lets us say "expend an average of at least 5 us of useful computation on each message in order to keep the overall cost of message passing below 10%".
Does 3M+ messages a second over tcp/loopback in my test.
(The fact that this go is competitive is pretty sweet.)
Also, this is only testing the time it takes the go client to write the messages to the socket, not the time the server takes to process the messages. So the benchmark would be the same with a noop server that reads and discards all incoming traffic. Am I wrong?
2 million messages per second is a pretty respectable number, and fairly comparable to systems like distributed Erlang and Akka.
http://shootout.alioth.debian.org/u64q/performance.php?test=...
From the technical side - routing is done using regular expressions, the speed of internal routing is proportional to number of different subscriptions.
Additionally, NATS have some interesting specific features - for example: if I remember correctly messages for slow consumers are just dropped.
NATS was originally a monolithic server, but I see that some work on clustering had been done [2].
[1] https://github.com/derekcollison/nats [2] https://github.com/derekcollison/nats/wiki/Cluster-Design