Disruptor-rs: better latency and throughput than crossbeam
github.com
github.com
[0] https://github.com/LMAX-Exchange/disruptor/wiki/Blogs-And-Ar...
Disruptor, Thompson, Farley, et al 2011
https://lmax-exchange.github.io/disruptor/files/Disruptor-1....
For latency sensitive application, it tends to have different purpose which mainly trade off CPU and RAM usage for higher low latency ( first ) and throughput later.
Actually looking at the code a bit, it seems like you could replace the select statements with the various handlers, and hook up some threads to them. It would indeed cook your CPU but that's ok for certain use cases.
But there's a bunch of stuff that isn't that part of the trading system, though. All the code that deals with the format of the incoming exchange might still be useful somehow. All the internal messages as well might just have the same format. The logic of putting events on some sort of queue for some other worker (task/thread) to do seems pretty similar to me. You are just handling the messages immediately rather than waking up a thread for it, and that seems to be the tradeoff.
Originally computers were expensive, and lots of users wanted to share a system, so a lot of OS thought went into this, LMAX flips the script on this, computers are cheap, and you want the computer doing one thing as fast as possible, which isn't a good fit for modern OS's that have been designed around the exact opposite idea. This is also why bare metal is many times faster than VMs in practice, because you aren't sharing someone else's computer with a bunch of other programs polluting the cache.
The cost is not high, it's much less expensive to have a CPU operating more efficiently than not processing anything because its syncing caches / context switching to handle an interrupt.
These libraries are for busy systems, not systems waiting 30 minutes for the next request to come in.
Basically, in an under utilized system most of the time you poll there is nothing wasting CPU for the poll, in an high throughput system when you poll there is almost ALWAYS data ready to be read, so interrupts are less efficient when utilization is high.
Of course you can get different numbers if you invent a nonstandard definition of utilization.
With 300W to 400W TDP (Xeon Sapphire 9200) and two CPUs per typical 2U case, cooling is a real challenge, hence my mention of water cooling.
Tokio and other libraries such as pthread allows thread to wait for something and wake up that particular thread when the event occurs. This is what allows scheduler to schedule many tasks to a very small set of cores without running useless instructions checking for status.
For foundational library, I think you want things that are composable, and low latency stuff are not that composable IMO.
Not saying that they are bad, but low latency is something that requires global effort in your system, and using such library without being aware of these limitations will likely cause more harm than good.
What you're looking for is io_uring on Linux or IOCP on Windows, I don't think osx has something similar, maybe kqueue.
Rust is not magic and you can compile both with llvm (clang++).
If you specify that the pointers don’t alias, and don’t use any language sugar that adds overhead on either side, the performance will be very similar.
The Rust implementation even needs to use a few unsafe blocks (to work with UnsafeCells internally) but is mostly safe code. Other than that you can achieve the same in C++. But I think the real benefit is that you can write the rest of your code in safe Rust.
* it's a Cargo package, which is trivial to add to a project. Pure-Rust projects are easier to build cross-platform.
* It exports a safe Rust interface. It has configurable levels of thread safety, which are protected from misuse at compile time.
The point isn't that C++ can match performance, but that you don't have to use C++, and still get the performance, plus other niceties.
This is "is there anything specific to C++ that assembly can't match in performance?" one step removed.
A more nebulous Rust perf thing is ability rely on the compiler to check lifetimes and immutability/exclusivity of pointers. This allows using fine-grained multithreading, even with 3rd party code, without the worry it's going to cause heisenbugs. It allows library APIs to work with temporary complex references that would be footguns otherwise (e.g. prefer string_view instead of string. Don't copy inputs defensively, because it's known they can't be mutated or freed even by a broken caller).
Standard C++ doesn't but `noalias` is available in basically every major compiler (including the more niche embedded toolchains).
And here’s one I saw linked on HN recently: https://github.com/0burak/imperial_hft/tree/main/distuptor