> Does 250k messages/sec at peak easily, More details in this presentation
https://www.youtube.com/watch?v=PB8dBnpaP8s250k messages/sec is pretty low.
It'd be useful (not to mention entirely feasible) to be able to handle 5M/sec bursts on a single system and aggregate sustained 50M/s+ (10Gbps network).
The system should be able to establish total order (serialization) locally and causal order over the network.
I did some tests and found out hardware limits are somewhere between 20-50M messages per second (serialized) on current consumer grade X86. Practical implementation would of course be slower.
Of course you need to go binary logging at that point, adhere to cache line boundaries, etc. mechanical sympathy. Maybe even do usermode networking at the aggregator server.
Binary logging is a must, because even something like string formatting is simply way too slow, by an order of magnitude.
In my quick tests I found the string processing hit to be surprisingly high, 10-50x. C "sprintf" and C++ stringstream are atrociously slow. (Surprisingly, considering sprintf has a pretty complicated "bytecode" format specifier parsing loop, it was still significantly faster than stringstream implementation.)
Timestamps are another huge performance issue. Typical system calls for high precision (a few microseconds or better) timestamp take several microseconds each - using one of those will alone drop performance to 100-500k range. It's of course much faster to use RDTSC, but then you have the issue that different CPU sockets have different offsets (sometimes large) and possibly also some frequency difference between them. Also older CPUs don't have invariant TSC, so their speed changes when CPU frequency changes.
(Some word of warning about Windows QueryPerformanceCounter: When you test it on your development laptop, it appears fast, because it's using RDTSC behind the scenes. But when it runs on a NUMA server, it often changes behavior and becomes 20-100x slower, because Windows starts to use HPET instead. RDTSC takes maybe ~10 ns to execute and it's not a shared resource. HPET takes 1-2 microseconds to read and is a shared system resource, concurrent access from multiple cores will make it slower.)