Fragmentation, MTU, MSS clamping and tunnels (2014)
fredrikjj.wordpress.com
fredrikjj.wordpress.com
One of those was tunnelling the traffic elsewhere and was getting packets "from the network" with packet sizes up to 65k despite the MTU of the tunnel being set to 1400. The transmitting host was respecting the MTU. I had to force clamp the MSS to the PMTU in iptables to get things working again.
FWIW, from someone who used to work on exactly this; neither did most of us :)
The org had fired^Wtold most of the network engineers to either interview for dev positions or leave. That left the rest of us with only dev experience to have to learn DC networking from scratch (without any mentors) while also building the infra to automate it, because, well software engineers should be able to learn network engineering overnight right?
I'll just say this, most of the devs I worked with had never coded before in a production environment, didn't know how to test their code (failing tests that would somehow pass in ways you'd think these were contest entries), and didn't even understand basic concepts in programming languages (I'm talking things like references vs values...).
On top of this, the tooling and infrastructure was so bad, even if you wanted to be productive you'd be fighting things like your build environment just randomly breaking after a git pull and having to reinstall visual studio or reformatting windows to fix it.
From what I've heard from people I still know on the inside, things haven't changed much (I'm very grateful for $MSFT though...)
MSS clamping is going away with protocols like QUIC and HTTP/3 encrypting L4 headers/options. PMTUD remains the proper way to handle things, or something like PLPMTUD if you're not willing to assume IP. Some things create path MTU black holes, commonly firewalls but also APs. APs can be particularly troublesome because nowadays they commonly don't terminate the wireless encryption, the controller at the other end of the tunnel does, which means they can't inject an ICMP to let the client know that packet is too big. MSS saves the day (for TCP) but otherwise the AP can either drop or fragment at the encap layer transparent to the client. I've yet to see a system that delivers the fragmented tunnel packet so the controller can ICMP back that the packet is too big but it'd be the cleanest thing in that scenario.
Also if you're going to limit tunnel size arbitrarily beforehand probably best to make sure the resulting packet ends up being able to fit into 1428. A lot of products use a 1300 "inside" MTU for this (and similar) reasons. 1400 + auto IPsec calc will lead to trouble in this scenario.
Don't drop "inside" MTU below 1280 in IPv6 networks. IPv4 is much more tolerant with the standard being 576 but many supporting less.