Maybe you wanna write a dedicated OS for it? Interesting project but I can’t blame them for not doing it.
Maybe you wanna write a dedicated OS for it? Interesting project but I can’t blame them for not doing it.
The network card also has hardware limits in the BW that it can handle, its latency. It is connected with the CPU via PCI-e usually, which has also latency and bandwidths, etc.
All this go to the CPU, which has latencies and BW from the different caches and DRAM, etc.
So you should be able to model what's the theoretical maximum of request that the network can handle, and then the network interface, the PCI-e bus, etc. up to DRAM.
The amount that they can handle differs, so the bottleneck is going to be the slowest part of the chain.
For example, as an extremely simplified example, say you have a 100 GB/s network, connected to a network adapter that can handle 200GB/s, connected with PCI-e 3 to the CPU at 12GB/S, which is connected with DRAM at 200GB/s.
If each request has to receive or send 1 GB, then you can at most handle 12 req/s because that's all what your PCI-e bus can support.
If you are then delivering 1 reqs/s then either your "model" is wrong, or your app is poorly implemented.
If you are then delivering 11 req/s, then either your "model" is wrong, or your app is well implemented.
But if you are far away from your model, e.g., at 1 reqs/s, you can still validate your model, e.g., by using two PCI-e bus, which you then expect to be 2x as fast. Maybe your data about your PCI-e bw is incorrect, or you are not understanding something about how the packets get transfer, but the model guides you through the hardware bottlenecks.
The blog post lacks a "model", and focus on "what the software does" without ever putting it into the context of "what the hardware can do".
That is enough to allow you to compare whether software A is faster than software B, but if you are the fastest, it doesn't tell you how far can you go.
But hey, doing science[0] is hard, better not be scientific instead /s
[1] science as in the scientific method: model->hypothesis->test , improve model->iterate. In contrast to the "shoot gun", or like the blog author called it, "whack-a-mole" method: try many things, be grateful if one sticks, no ragrets. /s
OP has defined the problem as speeding up an HTTP server (libreactor based) on Linux. So that's a context we assume as a base, questions like "what can the hardware do without libreactor and without Linux" are not posed here.
If you don't know, find out, because maybe X is already as fast as it can be, and there is nothing to speed up.
Sure, the OP just looks around and sees that others are faster, and they want to be as fast as they are.
That's one way to go. But if all others are only 1% as fast as _they should be_, then...
- either you have fundamentally misunderstood the problem and the answer to "how fast can X be?" (maybe its not as fast as you thought for reasons worth learning)
- what everyone else is doing is not the right way to make X as fast as X can be
The value in having a model of your problem is not the model, but rather what you can learn from it.
You can optimize "what an application does", but if what it does is the wrong thing to do, that's not going to get you close to what the performance of that application should be.
I aim for 9 Gbps per NIC, but I still see people settling for 3 Gbps total as if that's "normal".
y'know - it might be enough
Googling stuff like "Amazon AWS hardware TCP TOE" doesn't reveal anything. So we can't assume that either.
I'm not sure about AWS, but in Azure it is called "Accelerated Networking" and it is available in most recent VM sizes that have 4 CPUs or more.
It enables direct hardware connectivity and all offload options. In my testing it reduces latency dramatically, with typical applications seeing a 5x faster small transactions. Similarly, you can get "wire speed" for single TCP streams without any special coding.