Why are ethernet jumbo frames 9000 bytes? (2018)
blog.dave.tf
blog.dave.tf
Jumbo frames are not intended¹ to be adopted on the internet, and likely never will be. You need end-to-end control of devices to make sure they work correctly (as the article points out, even PPPoE ⇒ 1492 has PMTU issues on the open internet, it's best practice to run PPPoE on 1508 so you get back to 1500…) But they are very much used inside of administrative domains, e.g. for SANs and cloud interconnects. It's visible when you are a cloud customer, but absolutely not on the open from e.g. your home or mobile internet.
[¹] by the current understanding; I have no idea if Internet2 was ever as bold as thinking it might be possible to roll out 9k on the Internet "1".
So the simplification is that in v6 the legacy mode was eliminated.
I saw a graph at a the time showing how many single-bit errors the checksum misses, and the graph was flat-ish up to around 8-10k and rose after that.
IE: A 1-bit Parity bit is symmetric, and will *NEVER* miss a single-bit error. 1-bit Parity can only miss 2-bit, 4-bit, 6-bit, 8-bit... errors.
I have an expectation for Ethernet's 32-bit CRC to have a similar property.
EDIT: I looked it up at the CRC Zoo. Ethernet is __NOT__ symmetric. https://users.ece.cmu.edu/~koopman/crc/c32/0x82608edb.txt
That means 1-bit, 3-bit, 5-bit (etc. etc.) errors can slip through.
I remember being curious about why the graph didn't look like a simple function of the packet length (linear or some other simple function), but I did not procrastinate by investigating the function and learning about why.
The page I linked above clearly demonstrates that Ethernet-CRC32 is Hamming-distance 2 at 524288-bits (aka: 65536 bytes) of input, or 64kB. IE: Its still *impossible* to have a 1-bit error escape undetected on 64kB frames (let alone 9000-byte jumbo frames).
I don't know where the 1-bit errors start to occur, but its well beyond 64kB (which is all the CRC-zoo tested for).
--------
Maybe you're misremembering your graph / data? There's also CRC-16, or even CRC-5 (aka: USB uses CRC5). Based on various analysis, its clear that CRC32 is sufficient for any typical Ethernet frame of any size under 64kB (and probably for many frames of size larger).
I can't find the thing I looked at. Another paper I found now writes that 64kB is far beyond the limit, though: ¨With Ethernet, the FCS computation uses a 32-bit cyclic redundancy check (CRC-32). CRC-32 error checking detects bit errors with a very high probability. But as frame size increases, the probability of undetected errors per frame may increase. Due to the nature of the CRC-32 algorithm, the probability of undetected errors is the same for frame sizes between 3007 and 91639 data bits (approximately 376 to 11455 bytes). Thus to maintain the same bit error rate accuracy as standard Ethernet, extended frame sizes should not exceed 11455 bytes." https://web.archive.org/web/20110807131142/staff.psc.edu/mat...
The thing I read at the time didn't agree that 11k and 9k have the same accuracy, though.
There are NICs which only support 4K frames. Those are also "jumbo frames".
It's convention to call anything from 1501 to 9216 as "jumbo" and anything up to 64k "superjumbo".
There are standards, they just aren't relevant. 802.11(n) specifies 7935 bytes (A-MSDU), 802.3(as) specifies "frame envelope" at 2000 bytes. (The latter is to have room for encapsulation headers to be able to deliver 1500 regardless of encap.)
IPv4 and IPv6 both support 65k packet lengths.
Is this part true? Aren’t the larger data center interconnects using 9000 bytes internally? Also, wouldn’t fat large distance links be using larger packets? I would have thought that a 4% wire overhead (and who knows how much CPU) would have pushed people away from 1500 for anything but last mile.
Path MTU detection is all sorts of fragile, so exposing a tcp maximum segment size implying a MTU above 1500 is asking for trouble; actually even just 1500 isn't always a good idea [1]. There are many networks where too big packets are dropped without notifications, and if the other end is also on a 9000 MTU LAN, then you can get stuck. sigh
On the host side, packetization is not that big of a deal; nics can help, but it's also just not that many packets. Where hosts tend to run out of cpu is when you're dealing with much smaller packets.
Larger packets could help routers of course. And larger packets would probably mean fewer acks, which would be nice too. But, it's unlikely to happen unless mtu probing is enabled in more places.
[1] The best is to send the lesser of (theirs - 28) or your actual value. Well best is relative, this will result in the most successful connections, at the expense of more packets for networks where people aren't insane.
1) Dropping all ICMP packets on the floor for "security" reasons, which means that basic diagnostics is now impossible, you get random issues with software that uses PING, and Path MTU Discovery is broken forever.
2) Leaving internal firewall ports with the default settings, which are intended for Internet-facing ports. For example, for "denied" incoming traffic from the Internet, the correct response is to silently drop the packet. Internally, the correct thing to do is to respond with a NACK to instantly close the connection. Without this, you spend days and days chasing down random and difficult-to-troubleshoot timeouts and weird 30-second delays all over the place.
As a random example, Windows RDP makes a HTTP call out to the Internet from the server to verify a CRL. This is safe and secure. Without this, bad certificates can't be blocked. Unfortunately, this happens in the "SYSTEM" context, which tends not get proxy settings applied to it, so it is often blocked, with a 30-second timeout. This failure is cached for 24 hours on the server, which causes a maddening delay when you connect to servers. Every day. Every server. But you can't reproduce it, because the second time it won't happen. (It also won't happen if anyone else connects to the server right before you.)
Years later, I still get angry remembering the snarky comments by the firewall guy saying that I'm just imagining things.
At high volume mtu bottlenecks, the routers dropping packets are likely to limit how many ICMPs they send; otherwise they'll run out of CPU. This wouldn't be terrible if ICMPs weren't so commonly dropped.
Some bottlenecked routers don't have globally routable IP addresses, and may not be able to send ICMP at all. This can happen inside long haul networks, where there may be tunneling that reduces the effective MTU. This seems less common today, but I dunno? One 'recent' advance is that some PPPoE networks use 'mini jumbo' frames, with an ethernet MTU of 1508, so that the PPP mtu is 1500; that may be happening on other layers with tunnels too.
PS: Even at CERN, where cooperation is supposedly tighter, JFs were only an experiment, and, AFAIR, not a very successful one.
It may sound simple but it was pretty hard for them to reproduce.
You are correct that intermediate boxes all take a perf hit for the extra packets, because routers and other boxes are usually packets per second bound at some point, rather than bandwidth.
This normally only happened when servers arrived with new NICs and no one thought to give us a heads up.
But when it comes to external L3 peering between 3rd-parties (eg: the Internet), it's very rare that people will mess with the IP MTU. Unless you can guarantee that the IP PMTU is higher than 1500 bytes (across multiple segments that you may not control, including ISP last-mile) there is simply no benefit.
The reason the Internet still runs exclusively on 1500 byte IP MTU is that any lower MTU in the path will effectively make the jumbo segments useless. This means any PMTU-D problems customers experience would have been unnecessary and avoidable by just using 1500 everywhere.
Many public IX's (peering exchanges) have a jumbo vlan but it's almost always separate from the standard one that only allows 1500 byte MTU.
> 9001 bytes was the absolute max they could give to customers and still be able to do packet shenanigans.
https://news.ycombinator.com/item?id=30838866
See post for more detailed context.
Or maybe it's transitive, 8K pages influenced NFS transfer size, or retrospective, NFS transfer sizes influenced page size (I don't know the history of Solaris page sizes or NFS protocols).
Unrelated but fun: Of all the enterprise gear in my home network lab a $25 TP-Link consumer switch from Amazon has the highest MTU at 15k jumbo. TL-SG108E
Is it? As far as I can tell, T-Mobile US gives me a 1500 MTU with my LTE modem, and 1416 on my cell phone.
In addition to the IP MTU you also have layer2 headers, almost always Ethernet these days. Ethernet can optionally have one or more 802.1Q headers, among other things.
Networks that use only one vendor sometimes just configure their internal links at the maximum supported by that vendor/platform. In multi-vendor networks this is problematic because each vendor has different maximums. In that case it is best to set an internal standard like 9100 IP MTU that is within every vendor's limits and leaves plenty of room for all the overhead (layer 2 and 3) listed above.
It's important to note that some protocols like OSPF require all participants on a segment to have the exact same IP MTU configuration.
EDIT: clarified 9000 is the standard customer jumbo IP MTU size, not standard customer IP MTU size which is of course 1500.
The reverse is true also! From TFA:
> There’s certainly a performance incentive to not fragment [8192-byte] NFS traffic.
Nowadays, most NFS traffic is likely over TCP. But back in the day it was not. So dropping 1 frame out of 6 that comprise a packet, meant that you had to retransmit all 6. BTW this is why DNS has lots of compression and limits on UDP packet size before [compliant] implementations switch to TCP. NFS has no such provisions.
Compare that to say, a 56K modem that can handle all of 4 1500 byte packets/second.
Also, not only time to retransmit matters, but the amount of power we expend per packet. I think it's ridiculous to have thousands of times more overhead than necessary, and this likely hurts a lot more than the rare retransmission.
Most NICs have various types of offloading now, especially if you get into 10+ GigE, so how much of a bottleneck is frame processing nowadays? (Either on the send or receive side.)
Now we're moving to 100Gb and soon 400-800Gb and anything helps up there. I'm waiting for someone to standardise some way to reassemble jumbo frames into large 1MB+ 'application packets' using custom headers and reassembly algorithms directly in NICs (hopefully something standard like P4) before DMA-dispatching them to user memory.
At high enough speeds even offloads prefer frames not be tiny just because they can. Even most data center 100G+ switching is not line rate below 256 byte packets and the only thing it's designed to do is transport packets as quickly between ports it can without really inspecting them. Also encapsulations/overlays make the header loss ratio even greater across network to network links in the path.
In the end though most places can do without jumbo without a noticeable impact. There can be buffer scheduling problems with ridiculously large frames as well.