Vector Packet Processing
netgate.com
netgate.com
VPP here seems to be a "user-mode network stack", as far as I can tell. I was kind of attracted to the title because I was hoping for SIMD / Vector compute maybe even GPUs, but that doesn't seem to be the case.
Still, a usermode network stack is apparently a must-have for any very-high performance network application. I've never needed it, but a lot of optimizers talk about how "slow" Linux networking is when you actually benchmark it.
Kernels typically cannot use vector instructions because if they did they would need to save and restore the vector register state when servicing interrupts. There is a very large performance cost to doing that.
Moving packet processing into userspace means adding latency, including TLB pressure, in order to do the context switch.
I imagine that we might get some innovation by allowing to configure the system such that the kernel owns the vector registers and userspace is not allowed to use them. If your primary interest in vector registers/instructions is packet processing, and you're doing that in kernelspace, you might not mind it if userspace can't use those registers.
This isn't the case because VPP polls the NIC from userspace and never enters the kernel. There are no context switches.
https://doc.dpdk.org/guides/linux_gsg/linux_drivers.html discusses it a bit.
If I understood the VPP design and implementation details correctly they try to reduce the amortised cache misses for a batch of packets by running all packets in a batch through each software pipeline stage before processing the next stage. This should result in very good average instruction cache hit rates and should also help with data cache hit rates because packet headers are small and can be prefetched while the forwarding data structures e.g. 1 million IPv4 prefixes and their next hops can be hard to fit into L2 data caches and won't fit into L1 data caches.
I assume a carefully tuned implementation can make further gains by dedicating cores to specific pipeline stages to keep the data caches hotter at the cost of copying processed packet headers their next stage in new (sub-)batches. The actual packet content is only relevant for a few operations like encryption/decryption and modern highend NICs have line rate crypto engines to help with IPsec or TLS.
- amortizing cache misses are you mentioned
- better use of out-of-order, superscalar processors: by processing multiple independent packets in parallel, the processor can fill more execution units
- enable the use of vector instructions (SSE/AVX, VMX etc): again, processing multiple independent packets in parallel means you can leverage SIMD. SIMD instructions are used pervasively in VPPWhile this is true, it definitely doesn't seem to stop people like Apple [1] from using SIMD extensions in the kernel anyways. On ARMv8, it's an extra 512 bytes (32 qword registers) or eight (and change, depending on alignment) dirty cache lines. Whether or not this causes a serious performance impact will depend on how well the kernel can actually make use of SIMD (saving and restoring might cost 1% perf but if speeding up the kernel wins 5%, who cares!). It could be interesting to play with this on other kernels to verify these assumptions, it might be worth enabling in the kernel on some devices! Kernels do tend to do a lot of moving data around, which these extensions (on ARM, anyways) excel at.
[1] https://github.com/apple-oss-distributions/xnu/blob/main/osf...
You can read more about it here: https://s3-docs.fd.io/vpp/24.02/aboutvpp/scalar-vs-vector-pa...
EDIT: I guess from a SIMD-perspective, I'd have expected an interleaved set of packets, a-la struct-of-arrays rather than array-of-structs. But maybe that doesn't make sense for packet formats.
I can see how this approach described is "vector-like", even if the vector is this... imaginary unit that's parallelizing over the branch predictor instead of an explicit SIMD-code.
This "vector" organization probably has 99.999%+ branch prediction or something, effectively parallelizing the concept. But not in the SIMD-way. So still useful, but not what I thought originally based on the title.
(also, x86-64 CPUs with good gather instructions are rare, and sibling comments show that this is aimed at lower end CPUs. That makes SIMD even less relevant.)
And VPP is targeting high-end system and uses plenty of AVX512 (we demonstrated 1TBps of IPsec traffic on Intel Icelake for example). It's just very scalable to both small and big systems.
With memif especially, it's fast as heck. But you need to rebuild your apps to target memif. There's some pretty good drop in stdlib replacements for languages like Go, but it's still some work to use the DMA accelerated shared memory packet processing high speed userland mode that VPP is capable of. Ex: https://github.com/KusakabeShi/wireguard-go-vpp
Vpp has very good documentation: https://s3-docs.fd.io/vpp/24.02/ A very cool unique feature is the graph representation for packet processing, and the ability to insert processing nodes to the graph dynamically per interface at some point in the processing using features (https://s3-docs.fd.io/vpp/24.02/developer/corearchitecture/f...)
The same core will do 14.99Gbps of IPsec (aes-128-gcm, 1480 byte packets) using VPP, largely because it supports (VEX-encoded) VAES.
While these aren't ARM cortex A72s, they're quite close (cheap low power) for Intel.
Depends on the application, batching policy, compute intensity, etc. But you can put 8 NICs and 8 GPUs in one node (and have them communicate through nvlink, so huge intergpu bandwidth!) which I can't for CPUs. You can maybe also get some unobtainium A100X or CX7+H100 to skimp on PCIe if you're well funded...
My thinking is that one needs to have a set of packets before being able to start processing. The first packets arriving must wait until enough packets has arrived to fill min size of the vector. And if the last packet comes "late", the arrival time of the last packet adds to the time for the other packets, thus adding something that looks like variance.
I assume there are parameters setting min number of packets in a vector, and timeouts for when to accept packets into a given vector.
... And doesn't this also adds, creates a relationship between otherwise independent packets, potentially creating a way to tag packets through a network. Basically if I can control the arrival time of my packets to a router (I send them at a baseline fixed rate, but delay the transmit time with a pattern), packets that are then bunched together to be vector processed in the router will also be affected by this delay. I could possibly then observe this pattern at other places in the network. Thus tracing packets.
Possibly.
I wouldn't be surprised if this actually reduces variance.
For Linux user space solutions without kernel bypass if the above mentioned network accelerators are not installed (devices in customer's premise, etc), the recommended way is to use Netmap since it enabled direct access to the network interface card (NIC) buffers from user space otherwise you are at the mercy of Linux own notorious sk_buff [1].
Another alternative perhaps for more efficient buffering in Linux is to use PF_ring or the the-kid-on-block IO_ring but not sure they are being currently being utilized in VPP or not.
For good introduction on Linux Networking acceleration technology, this presentation is a good start [2].
[1] VPP docs: Create netmap:
https://docs.fd.io/vpp/17.04/clicmd_src_vnet_devices_netmap....
[2] Linux Networking: The meaning of acronyms eBPF, DPDK, XDP, VPP [video]:
> With experimental technologies, Linux has been shown to make some gains in artificial benchmarks, such as dropping all received packets
Is this a joke? A jab?
dropping a packet after processing overhead and/or explicit classification work is a useful benchmark yes
Cisco is the one who wrote it and open sourced it. Netgate is just putting a wrapper around other people's (Cisco's) code.
>Is this a joke? A jab?
No, dropping packets is step 1 to proving you can get past some of the current CPU bottlenecks. Actually doing something useful is obviously significantly more work, but no point bothering with that work if the CPU is still the bottleneck.
Cisco open sourced VPP in 2016 and we've been busy working on it ever since.
https://www.stackalytics.io/unaffiliated?module=github.com/f...