SOCKMAP – TCP splicing of the future
blog.cloudflare.com
blog.cloudflare.com
I love everything about it!
<3
Good, wholesome, internet collaboration right here.....
The other solutions definitely all would block if the send buffer is full.
Apart from that question it was interesting to learn that io_submit actually works on sockets too. I definitely need to read more about this one.
- backpressure
- what if both userspace and SOCKMAP are doing read() on a socket?
- can one use SOCKMAP to splice only a selected amount of data, and pick up the rest with read()
- What are the parser/verdict program semantics. What can they do.
- what is the SK_MSG abstraction, and how to benefit from it.
and more... some of this is discussed:
http://vger.kernel.org/lpc_net2018_talks/ktls_bpf_paper.pdf
On io_submit, we wrote about it here:
https://blog.cloudflare.com/io_submit-the-epoll-alternative-...
TLDR: it allows for batching, and with IOCB_CMD_POLL can be used as epoll alternative.
I guess the fact that one has to guess and try how those APIs behave makes me most nervous and would make me prevent from using those. An async IO system should be very well-behaved, and not block at arbitrary points.
But BPF is no that versatile, but at least it has a stable ABI.
https://blog.cloudflare.com/kernel-bypass/
https://blog.cloudflare.com/single-rx-queue-kernel-bypass-wi...
Also using dpdk and netmap kills usual tooling (from basics like tcpdump). We very much like iptables, xdp, conntrack, syn cookies and other technologies deeply embedded into linux kernel. Doing DDoS once on linux kernel is hard enough. We don't want to redo the logic for each of the possible kernel bypasses technologies:
https://blog.cloudflare.com/why-we-use-the-linux-kernels-tcp...
https://blog.cloudflare.com/syn-packet-handling-in-the-wild/
I don't see any fundamental reason why SOCKMAP wouldn't be fastest. It's just a matter of putting an effort into it.
For example, we noticed the poor splice(2) performance is due to a spinlock. We need to figure out just why this happens, but it doesn't seem like a major problem. Probably just a trivial regression.
Then you can get a view of every clock cycle and what is going on in the CPU, the caches, the hardware peripherals, the DMA controller etc.
You can start from a goal like "I want a TCP packet forwarded in under 5k clock cycles" and keep hacking code till you've met it.
Other software based tracing tools don't tend to be as good, and frequently rely on statistical techniques which hide what is really going on.
while data:
data = read(sd, 4096)
write(sd, data)
Not if that’s the write syscall, since it can return with fewer bytes than requested having been written. The full C code detects this case and crashes, which looks totally wrong.The Python-ish example should, at the very least, use sendall. But even that is potentially suboptimal.
https://blog.cloudflare.com/why-we-use-the-linux-kernels-tcp...