Watch Where the Linux Kernel Drops a Packet
linux.die.net
linux.die.net
You can also achieve mostly the same effect with perf via the skb:kfree_skb tracepoint. This has less performance impact, better post-processing options, and doesn't require the dropwatch specific kernel config option (NET_DROP_MONITOR).
this is a very well known phenomenon, permeating everything, everything around us. it's how advertising works.
Is anyone aware of an overview showing how to approach and get started with this tool?
I've found it useful to perform a baseline test to see "normal" drops and their rates, then run intense test traffic and compare.
You pretty much have to understand each reported call site by looking at the kernel source code. As you gain experience with the kernel network code, you learn which locations do what, and how to access the regular statistics for them. Having the (outdated but) relevant kernel networking books on your desk helps.
Wow, this changes everything. Thanks for clarifying. With this insight `dropwatch` sounds much harder to use then suggested.
https://github.com/torvalds/linux/blob/master/net/core/drop_...
I've recently made it on-by-default in NixOS and also collected some info about which other distros already have it on-by-default and since when: https://github.com/NixOS/nixpkgs/pull/85119
dropwatch is very useful, I used it to debug interruptions of our VPN setup between servers.
For example: perf record -e skb:kfree_skb
Thanks @2bluesc for creating the AUR for Arch (Yes, Arch + Manjaro user ;-)
NOTE: what surprised me was that Fedora 32 has dropwatch in its official repo (so simply dnf install dropwatch worked without fuss). It all made sense when discovering the author works for Red Hat when looking at the GitHub repo, well done.
especially the reason for the drop? what does +48 in "tcp_v4_rcv+48" mean, thanks and sorry if this was documented? and I TL;dr
The tcp_v4_rcv function has two obvious kfree_skb calls, here https://elixir.bootlin.com/linux/v5.6.13/source/net/ipv4/tcp... and a few lines below after "discard_it:" (there may be more if there are #define's or inlined function calls), and without your kernel image and debug info I cannot tell which one the offset corresponds to. Also, clean-up code like kfree_skb is often after a label ("out:") referenced by multiple goto's, and you cannot tell which goto was taken. However, often the function return value contains an error code that identifies it, and you can (often) grab that with perf by attaching a dynamic kprobe to the function exit (it's much easier than it sounds). Or attach a gdb to the kernel (easiest is if the kernel is in a qemu VM) and put a breakpoint on tcp_v4_rcv. There's also the inverse problem of "who called tcp_v4_rcv". Either gdb, or perf record -g, can tell you the stacktrace. (perf is less invasive, so better in production)
As an example, take https://elixir.bootlin.com/linux/v5.6.13/source/net/ipv4/tcp... : if the packet is a SYN belonging to a new connection, but has a bad TCP checksum, goto csum_error, and from there fall-through to discard_it: kfree_skb(skb).
This may sound laborious, and it is, but note that often you don't need to go to this effort. To me, as a troubleshooter, the precise reason might not be that relevant. The function name already tells me this packet has gone up into the TCP receive stack, which (basically) rules out entire problem areas like bridging and routing, tells me if this specific drop is even relevant for me, and/or lets me decide which simpler tools to use next.