How to drop 10M packets per second
blog.cloudflare.com
blog.cloudflare.com
I thought DPDK and friends don't use interrupts and instead map the buffer areas (for ring buffers etc.) used by the NIC into their process address space, directly polling those structures.
What confuses me is that deep in my memory something tells that DPDK had some kind of "deterministic mode" with guaranteed "zero drop at below n packets per n nanoseconds" that is driven by ring buffer "rollover sense" interrupts.
Or may be it just uses those to set the poll rate.
There is really no other option for line rate 10G packet handling in software. Linux can handle maybe 1.5 Mpps per core. That’s barely line rate for 1G. Multiple cores speed that up, but for sequencing purposes packets are usually distributed to different cores by hashing the IP 5-tuple, which means that more cores don’t help for a single connection (or a DDOS like the one in the article).
I’m going to switch my home router to use it, it seems to support PPPoE: https://docs.fd.io/vpp/18.07/clicmd_src_plugins_pppoe.html
And works with FRR: https://github.com/FRRouting/frr/wiki/Alternate-forwarding-p...
More docs on VPP can be found here, including installation, getting started, configuration and FAQ: https://wiki.fd.io/view/VPP
There’s even a Kubernetes networking plugin for VPP: https://github.com/contiv/vpp
At what packet size? At 1500B, 10Gbps is only 800Kpps.
In my experience working a large network, 300-600B seems to be the average. Worst case, that's ~420Kpps.
DPDK is an example of another option.
oh yes, vpp is builds on what is provided by dpdk.
But I think you might be right in this case, it seems more likely that the XDP processing is being done by the network card driver.
I'm curious because this would be the difference between "scale with more cores" or "scale with more NICs".
In terms of scaling with multiple CPUs, generally these days the NIC will be able to manage multiple separate ring buffers. Ring buffers can be dedicated to specific CPUs, and the NIC can distribute packets among the ring buffers (in hardware) by hashing on the packet's IP 5-tuple.[2] So long as you have packets that are distributed evenly by that mechanism (not the case when, e.g. the packets all belong to the same connection), you can scale with multiple CPUs. (Presumably the NIC can perform the distribution process in hardware at line rate.)
[1] At high packet rates, there might be an interrupt less frequently than for every packet.
[2] See section 7.1.8 of the linked manual.
Please correct me if I'm wrong.
Let me show how to render 100B polygons per second on a single cpu by using a gpu to do all the work... Argh.
because as you clearly know/demonstrate, the web can't be trusted by default