PacketMill: Toward per-Core 100-Gbps networking (2021)
dl.acm.org
dl.acm.org
My understanding is that there's some contention from the locking involved though which can make this not scale well, but maybe that could be avoided (e.g., maybe by trying to have per core buffers used by the kernel, although I'm not sure you'd have any control over the kernel threads).
You could go the DPDK route but that has downsides of your application needing exclusive access to the network interface. AFAIK, any network stack that supports multiple simultaneous applications typically involves at least 1 copy but I could be off.
Having worked with sk_bufs a bit I hope a copy isn’t needed for a network send. All the tooling I worked with around those was zero copy.
Never did it, but I guess you could reroute traffic you don't care about to the Linux network stack for regular routing.
https://doc.dpdk.org/guides-16.07/howto/flow_bifurcation.htm...
Not sure that's a realistic scenario though. You rarely expose a DPDK app directly on a regular / open subnet. You often have a dedicated app, with a dedicated NIC, on a dedicated subnet, for a dedicated traffic.
I was thinking of a situation where you want to also limit the system call overhead per send. (think maybe there's some number of 64k-ish chunks that you want to send to different clients, so each connection would involve multiple write per frame etc.,)
(That is, if you get my weird train analogy, and if I'm understanding your idea correctly.)
So instead of:
for each player:
for each block player needs:
# Write block to player's connection
You could do: for each player:
for each block player needs:
# Check if the block is already mapped in a kernel buffer. If it is, just append a write for it to the io ring.
# Otherwise, copy the block into a new buffer and send it.You don't have control from the application, but if you're really trying to get the most performance from your networking heavy application, you want to set up your server where the NIC rx and tx queues are cpu pinned, and the application thread that processes any given socket is pinned to the same cpu that the kernel is using for that socket. Then there's no cross cpu traffic for that socket (if all else goes well).
Hopefully all the memory used is NUMA local for the CPU and the NIC as well. That gets trickier with multiple socket systems, although some NICs do have provisions to connect to two PCIe slots so they can be NUMA local to two sockets.
(I mean visible like a videoconferencing application as opposed to something like invisible multicast dns)
Well multicast is not possible on internet so I guess that's why you don't see it often.
In private network centric protocols though (e.g. finance), multicast is ubiquitous.
And we are getting ever increasingly much much better about p2p dma- the future where the network card sends direct to the GPU which replies direct to the host, without ever bouncing through the cpu or it's caches, is in reach!
Very good article though. I'm an addict user of DPDK but more for latency rather than throughput, and it's interesting to see the challenges involved there.
Part of the problem I think is that they have to compare to other existing solutions (click, VPP etc) otherwise everybody will ask, but at the same time it is very difficult to do a fair comparison.
I'm a VPP developer hence I'm both biased and a knowledge-domain expert, but focusing on what I know, which is VPP: figure 11b, they compare VPP to PacketMill for some L2 patch workload. They claim their approach is fare because they use automated tooling to benchmark, good and we also do for VPP, but our results don't necessarily match theirs - and, surprisingly, our results are higher than what they claim for VPP.
PacketMill paper figure 11b for VPP for L2 patch for 64-bytes packets at 1.2GHz using DPDK MLX5 driver [1]: ~5Gbps
VPP for similar configuration [2]: ~7.5Gbps (already a 50% error margin?)
VPP using a more optimized NIC driver (native AVF vs DPDK MLX5) [3]: ~18Gbps (almost 2x what they seem to claim for PacketMill...)
All this to say comparing different solutions is terribly hard, and I'm not sure of the value of this kind of benchmarks.
For VPP we built CSIT [4] which is opensource under the Linux Foundation Networking to automate tests in the same environments, to compare between VPP and DPDK between releases and platforms.
[1] https://packetmill.io/docs/packetmill-asplos21.pdf
[2] http://csit.fd.io/report/#eNp1kd0OwiAMhZ8Gb0zN6MRdeaHuPQxidS...
[3] http://csit.fd.io/report/#eNp1kd0OgjAMhZ9m3pgaVkS88ULlPcwcVU...
IOW, one of their strong claim is that PacketMill is innovative because it avoids copying/converting uneeded metadata, but VPP is already doing that since years.
Finally, their claim to break the 100Gbps on single core @2.3GHz is cute, but again I'm afraid they're late to the party. They claim 12-13Mpps per core for 64-bytes packets for example but VPP can achieve 20+Mpps per core already for L3 forwarding (routing).
Again, benchmarking is hard, but I keep reading there claims over and over in academic papers when they're factually wrong for area I know about. I can only imagine what is happening for area I don't know :(