IPv6 Fragmentation Loss
potaroo.net
potaroo.net
The result? Most DNS queries went through. All chat applications worked. SSH generally worked, except when you started a full-screen terminal application. Smaller web pages load. Larger web pages loaded only partially.
Imagine non-technical users explaining their issue. "The Internet is half-broken. Please help."
God only knows how much hair I would have today if the world had figured MTU and fragmentation properly.
Blocked or dropped ICMP has caused me heartburn as well. I am pretty sure my co-workers are used to my "blocking ICMP is Evil Bad and Wrong" rants by now.
https://docs.aws.amazon.com/vpc/latest/userguide/vpc-network...
If your instances communicate with peers over the internet. So by default, AWS was built to assume your instances do not communicate over the internet, lol.
With usual TCP and UDP, you block by default, allow outgoing, allow replies and your internet works (consumer defaults). If you want, you allow specific incoming ports and it's good enough for most use cases. It's almost trivial to configure a reasonable basic firewall, and you can learn all you need in 5-10 minutes.
With ICMP, there is this long list of types and subtypes, some of which sound dangerous while others necessary. Which ones should you block and which ones should you allow?
Do you expect everyone to create an accurate threat model? People have lives. So, some people end up blocking ICMP and internet "works" except .. of course it doesn't.
This is the internet eqivalent of dumping your sewage in the lake because it's too difficult to deal with properly.
People that do this externalise the costs of their choices onto everyone else, and it takes a bunch of specialists to identify what's happening and educate the problematic network operator, and/or clean up the mess.
Doing a search on should I block ICMP, or what ICMP to block, answers are:
[1] No!!; some security issues; a lot of ICMP should be blocked; suggests further research
[2] you should selectively filter; example iptables rule to allow echo; assess evaluate and make your own rules
[3] listing of types and RFC recommendations for transit and local traffic
So I guess I should take "Should Be Dropped" and "Policy Should be Defined" from [3] and plug them into [2]. Why is it so hard?
Why isn't the answer: "No, defaults are safe, no need to block anything", or "Select one of: Endpoint / Site firewall / Internet router; customize if needed"?
IMO, this is a mess. This is the ultimate cause of all the sewage, time spent both debugging, and even more on learning what and how to block, configuring it all. Issues like this even hinder IPv6 adoption, because who is going to deal with all the complexity.
[1] http://shouldiblockicmp.com/
[2] https://blog.securityevaluators.com/icmp-the-good-the-bad-an...
[3] https://serverfault.com/questions/981558/which-ipv4-6-icmp-t...
> Doing a search on ... answers are:
Sorry, "Google search results" are not "up the chain". That's your first mistake.
There's a tonne of misinformation in Google search results. It's real hard to identify what's good and what's not if you're not a specialist in the field, so it's easy to fall victim to this and just believe well meaning well written blog posts.
Don't do this. When you don't know something, speak to someone who does know about the subject matter at hand. Not anonymous people on the internet, but someone in reality that you can have a conversation with. If it's a topic that's vaguely within the remit of someone you work with, that's a good place to start.
I have centurylink DSL with PPPoE and the thing that really bugs me is if your modem lost the PPPoE password, it can login with default credentials and ask for the passsword. CenturyLink clearly knows who I am without needing PPPoE authentication, so what do they get out of it?
Also what system even accepts ICMP based route redirects by default from the open internet?
Why not a higher or lower level protocol? Because this is the lowest level protocol which is end-to-end and has a path MTU. Lower level protocols would need to somehow handle the MTU and higher level protocols would need to fit their segments within the MTU.
Why 4096? Because this is a common page size on computers. Multiple of this would likely bring little benefits, due to GRO/GSO.
Also, it's my utopic world, so if you don't like it create your own. :))
Also, 4096 bytes at what level? Layer 2? And when you are using a tunnel, you add bytes and thus have to go higher than 4096 or go lower and reduce MTU, or fragment the packet to keep 4096 bytes.
You are just raising the limit, not solving the problem.
Don’t fixate this kind of policy in such a fundamental protocol.
ICMP is rate-limited, and 1500 byte packets work, so PMTUD doesn't kick. Hence, the connections hang if you hit the magic size.
Had to decrease MTU. Unfortunately, that doesn't solve remote-side UDP.
Worse, is that if a tunnel doesn't reduce the MTU and its packets are dropped by a router further down the line the the ICMP TOO_FAT response back to the tunnel endpoint not the original host. The tunnel endpoint has zero clue what to do with it and drops it on the ground, leaving the original host in the dark.
Even worse is when overzealous firewall admins start locking down everything they don't understand and that includes ICMP. "Some hackers could ping our networks!" Then you're truly up a creek.
Luckily MTU breakage is fairly easy to spot once you know what to look for. It's easy enough to fix too if your local firewall admins aren't the block everything type. Just fix the MTU on the tunnel endpoint and suddenly everything will start working.
One final note: If as a router or application developer you ever find yourself having to fragment packets, please, for the love of all that is holy, fragment them into two roughly equal sized packets. Don't create one huge packet and one runt from the leftovers. It hurts me to see packets coming in from some heavily tunneled source with 1400 byte initial packet followed by 3 or 4 tiny fragments. Or worse, 3 or 4 tiny out of order fragments followed by the 1400 byte bulk because the MAC prioritized small packets in its queue and send them first.
Ideally, you should not want to fragment IP packets. It is far better to do MSS clamping in TCP to prevent the size of the TCP payload to grew above the size of the max MTU in the path. TCP can stitch together the data at the other node just fine, without routers in between having to fragment the IP packets, which will kill your bandwith in comparison.
One thing I like about IPv6 is the minimum MTU of 1280. That's big enough that if I'm really uncertain about the environment I can just set my MTU to that and avoid future headaches without impacting performance too badly. IPv4's 576 minimum nearly triples the number of packets you generate, which is really noticeable when you're running routers close to their limit. Forwarding speed is often dominated by packets per second, not bytes per second.
IPv4 minimum MTU is... 68 bytes.
>Every internet module must be able to forward a datagram of 68 octets without further fragmentation.
Apparently, their customers were tunneling IP packets through another protocol, meaning that instead of sending an IP packet in a Ethernet frame, they were sending an IP packet in an X packet in an Ethernet frame. Since, like IP and Ethernet packets, the X packets need to contain some information related to the protocol, there was less room for the IP packet. I.e. the MTU was lower.
When you set the MTU on your Operating System (OS), it refers to something slightly different. Instead of "this is the maximum size packet that will fit", it means "assume this is the maximum size packet that will fit". You can use that setting to force your OS to send smaller packets if you know the MTU is lower than your OS thinks.
MTU doesn't refer to anything different in that case, MTU always refers to the maximum size frame "this" node can handle. It also doesn't always mean the OS assumes it's the maximum size packet that will fit on a path it's the maximum size Ethernet frame the OS knows will fit on the NIC. The OS has other methods for assuming things about a path. Forcing MTU lower does force the OS to assume any path is never more than the MTU though which is why it works as a fix.
A typical symptom is that things (e.g., ssh) hang. This typically happens when you use a protocol that uses Kerberos in an Active Directory domain with very large Kerberos tickets. PQ crypto would do this too.
http://icmpcheckv6.popcount.org/
(v4 version http://icmpcheck.popcount.org/ )
it answers:
- can fragments reach you
- can PTB ICMP reach you
hope it's useful. Prose: https://blog.cloudflare.com/ip-fragmentation-is-broken/
Notice: it's easy to run the tests headless with curl if you need to see if your server is configured fine.
Fun fact is that it's very much not easy to accept/send fragmented packets from linux. I learned the hard way what `IP_NODEFRAG` is about.
The lack of on-path fragmentation in IPv6 is definitely on purpose. It was a mistake in IPv4 and would be silly to replicate in IPv6. The fragmentation header in IPv6 is effectively useless. It can only be done at the endpoints, and if that's the case the application should be doing it, not the stack. Instead IPv6 mandates path MTU discovery, which is the correct solution.
For UDP, it's not so simple; IP fragmentation does allow for large data, all or nothing processing, without needing application level handling, but the cost of fragmentation is high.
The out of band signalling when sending packets that are too large is too easy to break, and too many systems are still not setup to probe for path mtu blackholes (the biggest one for me is Android), and the workarounds are meh, too.
Another option would be for IP fragments to have the protocol level header, so fragments could be grouped by the full 5-tuple (protocol, source ip, dest ip, source port, dest port) and kept if useful or dropped if not, without having to wait for all the fragments to appear.
You get the same issue putting the information on the fragments. Now there is no "IP layer" there is just "well we're using IP+UDP today and how that is right now should forever be baked into this hardware that will be here for 20 years" which is excatly the problem that led Google/IETF to push headers deeper with HTTP/3 to get out of that mess.
You also can get an in-band signal that you're being fragmented in the middle without changing IP. E.g. TCP already negotiates an MSS, if IP fragments at the start of a group come in smaller than that you know there is something fragmenting in the middle.
In the middle fragmentation is not really something that happens very often. IPv6 prohibited it, but in IPv4, nearly all packets are marked do not fragment, because IP routers weren't fragmenting much anyway; I think it's more likely to get an ICMP needs fragmentation packet on a too big packet with Don't Fragment, than to actually get fragmented delivery.
Also, MSS is mostly not a negotiation; most stacks send what they think they can receive in the syn and the syn+ack. The only popular stack that sent min(received MSS, MSS from routing) was FreeBSD, but they changed to the common way in 12 IIRC; which in my opinion is a mistake, but I don't have enough data to show it... actually what seems best is to send back min (received MSS - X, MSS from routing), where X is 8 or 20, depending on if you more of your users are on misconfigured PPPoE or behind misconfigured IPIP tunnels.
The vast majority IPv4 traffic does not have the DNF bit set. Your logic of why they would doesn't even make sense as setting DNF only means it'll drop on routers that would fragment have fragmented not improve the situation with the ones that wouldn't have.
MSS is definitely a negotiation but a negotiation just between the TCP aware nodes not along the whole IP path which is why I say endpoint stacks can use it to detect if the IP path is fragmenting by comparing incoming IP fragment sizes to the MSS.
I don't think that's correct. Windows, Linux (including Android), FreeBSD, macOs and (Apple) iOS all default to sending Don't Frag. Together, they form the overwhelming majority of IPv4 traffic.
Nobody actually wants fragments, and very few fragments are seen in regular activity. When I ran high traffic servers, I would normally see a few fragments a minute per server, except for when we were under a chargen reflection DDoS attack; Microsoft Services for Unix sends back a hunormous UDP reply which is of course fragmented, and that caused some trouble with fragment reassembly buffers (dropping the buffer size to the minimum solved that, although during a DDoS the couple of people sending fragments would have a bad time; can't win them all).
I think sometimes there's fragments in large UDP DNS replies; but it's generally best to avoid large UDP DNS replies because they often get dropped by poor decisions in network design and software, and there's probably a better way to do what you need to do.
Stacks could use MSS and fragment size to do something cool, but they don't; in part, because nobody sends fragments anymore, and in part, because there's a lot of other cool stuff to do so path MTU gets left behind in a lot of places, like Android. :(
Edit to add:
> At that point it pretty much amounts to "send a message in a raw IP header to the destination rather than a message in an IP+ICMP header to the source".
The thing is, we know that the router that's in a position to truncate the packet can probably get a packet to the destination (otherwise, it wouldn't have this packet). We don't have anything to indicate that it can get a packet to the source; and it turns out that often it can't. An in-band indication would be so much more useful than an out of band (which doesn't work), or fragmentation (which people don't want). Yes, the indication goes to the wrong party, the destination will get the information, but the source needs it; but most interactions on IPv4 are two way, so presumably the destination can tell the source that it needs to send smaller packets, again in-band with the communication it already has, that presumably works.
Yeah no stacks do it that I know of, just a theoretical way one could without having to rely on in-path devices sending messages or dropping traffic. I think the v6 path of drop it and send it back was better though. If it's blocked that's their fault :p.
Regarding sending the "message too long" note via in band channel instead of its own ICMP message I actually think that's a great angle as it'll get through more FWs I'm just hung up on trying to do something the truncated payload once it arrives... but I guess that doesn't really matter from a protocol perspective - the important bit is the note on the MTU issue arrive in band and the OS can decide what to do with any remaining payload for inner protocols after that as it's not IP's problem. v6 already got rid of the IP checksum so the only other things that might need to be updated are NICs/FWs/NAT boxes that check TCP/UDP/protocol checksums to ignore them if the header had that note.
Another point to add to this is that many routers will rate limit control plane traffic (ICMP, BGP etc), or prioritize data plane traffic ahead of it, so the ICMP too large message is more likely to be lost than a hypothetical in-band indication. It's the same reason you sometimes see packet loss in the middle of a traceroute, but not the end.
Designing hardware for a single purpose most of the times is not a good idea. That means that you are stuck with the implementation that was implemented in the hardware for a very long period. Also, an hardware implementation can't take into account (and handle) all particular cases, for example malformed packets.
IP was designed to be flexible, you could have in theory used whatever L2 protocol you wanted, and whatever L4 protocol you wanted. Thanks to hardware implementation that considered only Ethernet as L2 protocol and TCP or UDP as possible L4 protocol anything was innovated.
The reason that we are so slow to adopt IPv6, same reason. What would have take to just adopt it if we didn't have hardware implementation of IPv4? Software update and you are done.
The problem is that in the world there are many not so good programmers that write inefficient code and thus people think that they need to implement things in specialized hardware to make them faster. You don't, at least in most cases.
Yes, yes it does. CPUs will never catch purpose built chips for networking.
Networking chips are getting more programmable, though. The most known effort is P4 (https://p4.org). Even for vendors that don't support P4, they are still internally making their hardware more fungible, which is good.
> The reason that we are so slow to adopt IPv6, same reason.
IPv6 has been slow to adopt because it solves a problem most people have not had, in an overly complicated way.
You can upgrade the control plane of a router with traffic still flowing thanks to the data plane being seperate. You can even make another router do control plane for the data plane in another router. Not to mention the massive performance benefits.
Have fun routing several Tb/s of traffic without dedicated ASICs....
so unless the connection didn't respond to anything at all after it failed to accept your password, your problem is not with the ssh tunnels or due to path/mtu issues themselves.
I know. And if it did you would see it in commands you type, too. Not just in invisible passwords.
I just said this mystery came to my mind when reading the story of being able to communicate with som web servers but not others. Not trying to claim we have that same MTU problem.
Actually now that I type it, I had the problem of half-broken internet myself in the past when roaming internationally with my phone. I needed to use a VPN to contact certain web servers. Need to keep this in mind when roaming becomes relevant again some day... (This happened in the UK and somewhere in the EU and affected e.g. my online banking and a Linux Foundation conference web site. I don't think censorship was the culprit.)
Where does that quoted portion come from. The mind of the OP author or someone else.
If IPv6 was just a 128-bit version of IPv4, I would be an IPv6 user.
As long as it continues to work, on the networks I control, I will prefer the relative simplicity of IPv6.
Relative to IPv4, IPv6 is more complex.
Running v6 is similar in complexity to v4, but a lot simpler than v4+NAT. Since NAT is more or less a necessity in v4 these days, v6 ends up being simpler in practice.
After similar problems with ssh, I routinely add "-4" to any ssh invocation. I just cannot be bothered with a protocol that adds nothing to my life but problems and headaches.
I'll deal with ipv6 when I hear that ipv4 is being deprecated. As in, most likely never.
Secondly, I don't get this mentality that IPv6 is some totally inferior thing and I really don't get people advising others to disable IPv6 just because they don't understand it.
Statements like "IPv4 does everything I need!" are, ultimately, totally missing the point. The fact is that we have hit the scale ceiling of the IPv4 address space and other cracks are showing and now the entire world needs something that will scale for the next few hundred billion devices.
IPv6 is the protocol that will lift the scale ceiling higher, not just for you and your needs, but for everyone. It won't really change anything performance-wise, nor will it change how higher-layer protocols like TCP and UDP work, but that's intentional.
You can hold out on principle if you like but it won't gain you anything. The world will just migrate around you eventually.