Yeah, receiving packets is fast when you aren't doing anything with them.
Yeah, receiving packets is fast when you aren't doing anything with them.
Forwarders are usually not doing much with packets, just reading a few fields and choosing an output port. These are very common function in network cores (telcos, datacenters, ...).
DPDK is not well-suited for endpoint application programming, even though here you can still squeeze some additional performance.
But don't dismiss a framework that is currently so widely deployed just because you are not familiar with its most common use-case.
For these types of applications there are proprietary implementations that you can buy from vendors that are more suited to latency sensitive applications.
The next level of optimization after kernel bypass is to build or buy FPGAs which implement the wire protocol + transport as an integrated circuit.
Funnily, one of the biggest DPDK feature is an API to program smartNICs exactly in that way.
I think the neatest thing that sticks out to me is how logging in implementations of low latency trading applications was essentially pushed to the network. Basically they just have copper or fiber taps between each node in their system and it gets sent to another box that aggregates the traffic and processes all the packets to provide trace level logging through their entire system. Even these have solutions you can buy from a number of vendors.
Supercomputing people have done this repeatedly and successfully.
A more general purpose stack is UDP-based applications.
I wouldn't do this with TCP, which I agree is complicated and difficult to get right.
You can do XDP as well and gain zero copy but for a small subset of NICs and you still need to implement you own network stack I don't see ac way to avoid it much. There are existing TCP stacks for DPDK that you can use as well.