Has hardware-offload (for AES et al) got faster still, or that keeping CPU busy in the data path for 100gbps workloads is not ideal, or something else? The WireGuard website claims ChaPoly is at least as fast as hardware-accelerated AES. And that it can be further sped up with SIMD.
https://www.wireguard.com/known-limitations
> now we’re on to the next order of magnitude together
Curious: Is this work currently in progress? If so, what's more that's still lined up? The previous GRO/GSO(/LRO, too?) improvements were incredibly impressive (even to u/majke, https://news.ycombinator.com/item?id=35567268).
Thanks.