How 1500 bytes became the MTU of the internet
blog.benjojo.co.uk
blog.benjojo.co.uk
The data isn't encoded in the usual ways, so even 4 hours of begging FFMPEG were to no avail.
A few glances at wireshark payloads, the roughly translated documentation, and weighing my options, I embarked on a harrowing journey to synthesize the correct incantation of bytes to get the device to give me what I needed.
I've never worked with RTP/RTSP prior to this -- and I was disheartened to see nodejs didn't have any nice libraries for them. Oh well, it's just udp when it comes down to it, right?
SO MY NAIVETE BEGOT A JOURNEY INTO THE DARKNESS. Being a bit of an unknown-unknown, this project did _not_ budget time for the effort this relatively impromptu initiative required. An element of sentimentality for the customer, and perhaps delusions of grandeur, I convinced myself I could just crunch it out in a few days.
A blur of coffee and 7 days straight crunch later, I built a daisy chain of crazy that achieved the goal I set out for. I read rfc3550 so many times I nearly have it committed to memory. The final task was to figure out how to forward the stream I had ensorcelled to another application. UDP seemed like the "right" choice, if I could preserve the heavy lifting I had accomplished to reassemble the frames of data.. MTU sizes are not big enough to accomodate this (hence probably why the device uses RTP, LOL.). OSX supports some hilariously massive MTU's (It's been a few days, but I want to say something like 13,000 bytes?) Still, I'd have to chunk and reasemble each frame into quarters. Having to write _additional_ client logic to handle drops and OOO and relying on OSX's embiggened MTU's when I wanted this to be relatively OS independent... and the SHIP OR DIE pressure from above made me do bad. At this point, I was so crunched out that the idea of writing reconnect logic and doing it with TCP was painful so I'm here to confess... I did bad...
The client application spawns a webserver, and the clients poll via HTTP at about 30HZ. Ahhh it's gross...
I'm basically adrift on a misery raft of my own manufacture. Maybe protobufs would be better? I've slept enough nights to take a melon baller to the bad parts..
CONTEXTLESS, HEADERLESS, ENDLESS BYTE STREAMS OF COURSE, where the literal, idealized (remember udp) position of each byte is part of a vector in a non euclidean coordinate system.
I would love to read a collaborative work between you and James Mickens -- this genre of writing seems sadly under-present in the computing world...
My agent is a tin can. I think she used to hold beans. Sometimes I put a few smashed nickels in her and rattle. While I do this, I pretend she's reading me my messages, and I'm like "oh no, I would never consent to a biopic directed by THAT charlatan." and then we laugh and laugh.
Oh how we laugh.
For non-native speakers: embiggened means huge, enlarged, overgrown.
I am not a native speaker of English either
The word was created as a joke in a Simpsons episode, a word used in Springfield only. It is described as "perfectly cromulent" by a Springfielder, which is evidently meant to mean "acceptable" or "ordinary" but is another Springfieldism.
The joke may be lost on future generations who don't realise they're not normal words.
The wiki page talks about getting 5% more data through at full saturation but it doesn’t mention an important detail that I recall from when it was proposed.
It turned out with gigabit Ethernet or higher that a single TCP connection cannot saturate the channel with an MTU of 1500 bytes. The bandwidth went up but the latency did not go down, and ACKs don’t arrive fast enough to keep the sender from getting throttled by the TCP windowing algorithm.
If I have a typical network with a bunch of machines on it nattering at each other, that might not sound so bad. But when I really just need to get one big file or stream from one machine to another, it becomes a problem.
So they settled on a multiple of 1500 bytes to avoid uneven packet fragmentation (if you get half packets every nth packet you lose that much throughput). Somehow that multiple became 6.
And then other people wanted bigger or smaller and I’m not quite sure how OS X ended up with 13000. You’re gonna get 8x1500 + 1000 there. Or worse, 9000 + 4000.
My very industrious teammate did 75% of the work (4 man team, I did 20%, if you are generous with the value of debugging). One of the things we/he tried was to just drop packets that arrived out of order rather than reorder them. Turned out the reordering logic was reducing framerates. So he ran some trials and looked at OOO traffic, and across the three or so routers between source and sink he never observed a single packet arriving out of order. So we just dropped them instead and got ourselves a few more frames per second.
I'm using a KoalaBarrel. Koalas receive envelopes full of eucalyptus leaves. Koalas have to eat their envelopes in order. First koala to get his full subscription becomes fat enough to crush all the koalas beneath him. Keep adding koalas. Disregard letters addressed to dead koalas.
The problem is that the world went wireless, so maximum link speeds grew a lot but minimum link speeds are still relatively low. A single 64kB packet tying up a link for multiple milliseconds—unconditionally delaying everything else in the queue by at least that much—is not what we want.
I would argue: the problem is that the MTU isn't negotiated at all, but especially not based on link availability.
At this point 1500 is the standard, we can’t ever hope to increase it without a way to negotiate the acceptable value across the entire transmission path - that’s what IPv6 gives us.
802.11 already includes a fragmentation and reassembly mechanism at the 802.11 level, distinct from any end-to-end IP fragmentation. Unlike IP fragmentation, fragments are retransmitted if lost. So you could use 802.11 fragmentation for large packets sent at slow link speeds to avoid tying up the link for a long time.
For every thing besides real-time maybe.
All Ethernet adapters since the first Alto card had self-clocking data recovery [1].
Clock accuracy was never a problem, as long as it was withing the acceptable range required for PLL lock/track loop.
The reason for 1500 MTU is that for packet-based systems, you don't want infinitely large packets. You want small packets. but large enough so that packet overhead is insignificant, which in engineering terms means less than 2%-5% overhead. Thus 1500 max packet size. Everything above that just makes switching and buffering needlessly expensive, SRAM was hella expensive back then. Still is today (in terms of silicon area).
Look at all the memory chips on Xerox Alto's Ethernet board (below) - memory chips were already taking ~50% of the board area!
[1] Schematic of the original Alto Ethernet card clock recovery circuit: https://www.righto.com/2017/11/fixing-ethernet-board-from-vi...
EDIT: Lol! Author has completely replaced erroneous explanation with correct explanation, including link to seminal paper about packet switching. Good.
Why does a larger MTU make switching more expensive?
And why does it effect buffering? Won't the internal buses and data buffers of networking chips be disconnected from the MTU? Surely they'll be buffering in much smaller chunks, maybe dictated by their SRAM/DRAM technology. Otherwise, when you consider the vast amount of 64B packets, buffering with 1500B granularity would be extremely expensive.
> Why does a larger MTU make switching more expensive?
Switching requires storage of the entire packet in SRAM.
Larger MTU = More SRAM chips
If existing MTU is already 95% network efficient (see paper), then larger MTU is simply wasted money.
Which raises another point in relation to the 1500 MTU - all of the CRC checks in various protocols were designed around that number. Even the checksum in the TCP header stops being effective with larger frames, so you end up having to do checksums at the application level if you care about end to end data integrity.
https://tools.ietf.org/html/draft-ietf-tcpm-anumita-tcp-stro...
"The advantage of this technique is speed; the disadvantage is that even frames with integrity problems are forwarded. Because of this disadvantage, cut-through switches were limited to specific positions within the network that required pure performance, and typically they were not tasked with performing extended functionality (core)." [2]
[1] https://en.wikipedia.org/wiki/Cut-through_switching
[2] http://www.pearsonitcertification.com/articles/article.aspx?...
Bitflips are very rare in a datacenter environment, typically caused a bad cable that you can just replace or clean. And crc check is done at the receiving system or router anyway.
Hmm. Why is this? It seems if we have a CRC-32 in Ethernet (and most other layer 2 protocols), we'll have a guarantee to reject certain types of defects entirely... But mostly we're relying on the fact that we'll have a 1 in 4B chance of accepting each bad frame. Having a bigger MTU means fewer frames to pass the same data, so it would seem to me we have a lower chance of accepting a bad frame per amount of end-user data passed.
TCP itself has a weak checksum at any length. The real risk is of hosts corrupting the frame between the actual CRCs in the link layer protocols. E.g. you receive frame, NIC sees it is good in its memory, then when DMA'd to bad host memory it is corrupted. TCP's sum is not great protection against this at any frame length.
> Which raises another point in relation to the 1500 MTU - all of the CRC checks in various protocols were designed around that number.
Now we have a new claim:
> The risk is that multiple bits in the same packet are flipped, which the CRC can’t detect
Yes, that's always the risk. It's not can't detect-- it almost certainly detects it. It's just that it's not guaranteed to detect it.
It has nothing to do with MTU-- Even a 1500 MTU is much larger than the 4 octet error burst a CRC-32 is guaranteed to detect. On the other hand, the errored packet only has a 1 in 4 billion chance of getting through.
> 100G Ethernet transmits a scary amount of bits so something that would have been rare in 10Base-T might happen every few minutes.
The question is, what's the errored frame rate. 100G ethernet links have error rates (in CRC errored packets per second) compared to the 10baseT networks I administered. I used to see a few errors per day. Now I see a dozen errors on a circuit that's been up for a year (and maybe some of those were when I was plugging it in). 1 in 4 billion of those you're going to let through incorrectly.
Keep in mind faster ethernet has set tougher bit error rate requirements and we have an undetected packet error time of something like the age of the universe if links are delivering the BER in the standard.
(Of course, there's plenty of chance for even those frames that get through cause no actual problem-- even though the TCP checksum is weak, it's still going to catch a big fraction of the remaining frames).
The bigger issue is that if there's any bad memory, etc, ... there's no L2 CRC protecting it most of the time. And a frame that is garbled by some kind of DMA, bus, RAM, problem while not protected by the L2 CRC has a decent risk of getting past the weak TCP checksum.
> EDIT: Lol! Author has completely replaced erroneous explanation with correct explanation, including link to seminal paper about packet switching. Good.
Don't be a jerk. Being right doesn't give you the right to make fun of people.
I am surprised the PLLs could not maintain the correct clocking signal, since the signal encodings for early ethernet were "self-clocking" [1,2,3] (so even if you transmitted all 0s or all 1s, you'd still see plenty of transitions on the wire).
Note that this is different from, for example, the color burst at the beginning of each line in color analog TV transmission [4]. It is also used to "train" a PLL, which is used to demodulate the color signal transmission. After the color burst is over, the PLL has nothing to synchronize to. But the 10base2/5/etc have a carrier throughout the entire transmission.
[1] https://en.wikipedia.org/wiki/Ethernet_physical_layer#Early_...
[2] https://en.wikipedia.org/wiki/10BASE2#Signal_encoding
The datasheet for the NS8391 has no such requirement for PLL sync.
https://archive.org/details/bitsavers_nationaldaDataCommunic...
Futher PLLs have not got a lot better, but a lot worse. Maybe back when 10BASE2 was introduced you could train a PLL on 16 transitions and then have acquired lock but there's no way you can do that anymore (at modern data rates). PCI express takes thousands of transitions to exit L0s->L0, which is all to allow for PLL lock.
My best guess for the 1500 number is that with a 200ppm clock difference between the sender and receiver (the maximum allowed by the spec, which says your clock must be +-100ppm) then after 1500 bytes you have slipped 0.3 bytes. You don't want to slip more than half a byte during a packet as it may result in duplicated or skipped byte in your system clock domain. (2001e-6)1500=0.3.
I figured this is what the interframe gap is for - to allow the FIFO to completely drain.
I left the lSP backbone and large enterprise WAN field around that time and can't speak to more recent technologies.
I feel like the last piece we’re missing in this story is the performance impact of fragmentation. Like why not just set all new hardware to an MTU of 9000 and wait ten years?
It used to have even more stuff, but I think he removed a lot when he got his book published.
The hardware in question is Ethernet NICs. However, for you to set the MTU on an Ethernet NIC to 9000, every device on the same Ethernet network (at least the same Ethernet VLAN), including all other NICs and switches, including ones which aren't connected yet, must also support and be configured for that MTU. And this also means you cannot use WiFi on that Ethernet network (since, at least last time I looked, WiFi cannot use a MTU that large).
Some routers will generate the ICMPs, but are rate limited, and the underlying poor configuration means that the rate limits are hit continously and most connections are effectively in a path mtu blackhole.
>Some routers will generate the ICMPs, but are rate limited, and the underlying poor configuration means that the rate limits are hit continously and most connections are effectively in a path mtu blackhole.
Sure. But I'm not about to sit here and name all the different reasons for folks. And since most here do not have a strong networking background running consumer grade routers at home, it seemed most applicable.
I could have used a more encompassing term like PMTU-D blackhole, but I didn't.
Now you need Path MTU discovery, which as the article indicates, has its own set of issues. (Overhead from trial and error, ICMP being blocked due to security concerns, etc...)
They blocked ICMP, do you deserve what you get?
Agreed
> and generally do
Agreed.
Now if you can make it 'will always just push packets', we'll be golden.
Unfortunately, there are enough ATM/MPLS/SONET/etc networks being run by people who no longer understand what they're doing, that we're never going to get there.
To make matters more entertaining, IPv6 depends on icmp6 even more.
You can have working fragmentation if you have two separate Ethernet segments, one for 1500 and the other for 9000, connected by an IP router; the cost (assuming no broken firewalls blocking the necessary ICMP packets, which sadly is still too common) is that the initial transmission will be resent since most modern IP stacks set the "don't fragment" bit (or don't include the extra header for IPv6 fragmentation).
Almost all IP packets on the internet at large have the 'do not fragment' flag set. IP defragmentation performance ranges from pretty bad to an easy DDoS vector, so a lot of high traffic hosts drop fragments without processing them.
If we had truncation (with a flag) instead of fragmentation, that might have been usable, because the endpoints could determine in-band the max size datagram and communicate it and use that; but that's not what we have.
Even if you really want devices to use JF, some fail miserably because it's just not well thought out.
Routers can fragment the packets, switches can't. So that would be pretty chaotic for non-techie installed equipment.
There are also certainly a shit load of them in closets and top-of-rack all over where I work.
Because a node with a MTU of 9000 will very likely be unable to determine the MTU of every link in it's path. At best, you'll see fragmentation. At worst, the node's packets will be registered as interface errors when it encounters an interface lower than 9k. Neither of those are desirable.
IEEE could define a way to support larger frames. 'just wait 10 years' doesn't strike me as the best solution, but at least it is a solution. In my opinion a better way be if all devices would report the max frame length they support. Bridges would just report the minimum over all ports on the same vlan. When there are legacy devices that don't report anything, just stay at 1500.
IETF can also do someything today by having hosts probe the effective max. frame length. There are drafts but they don't go anywhere because too few people care.
EDIT: As pointed out below, I failed to account for the clock-rate being 25% faster than the bit-rate in my original assertion that Ethernet over twisted-pair was only 80% efficient due to the encoding (see below)
Gigabit Ethernet is more complicated, and it uses multiple voltages and all four pairs of wires bidirectionally. So it is not just a single serial stream of on/off.
Actually, they already accommodated for this in the advertised speed.
In other words, a 1 GbE SerDes runs at 1.250 Gbit/s, so you end up with an actual 1 Gbit/s bandwidth.
The reason you don't actually hit 1 Gbit/s in practice is due to other overheads such as the interframe gaps, preambles, FCS, etc.
Actually Gigabit Ethernet is highly efficient; it can actually give you 98-99 % of line rate as the payload rate.
But I'm not really sure about the clock sync limitations being a factor here. It was way back in the deepest past.
What I do remember vividly is the mess that physical layer networking evolved into over the years thanks to dial-up and DSL (ever had to set your MTU to 1492 to accommodate an extra PPP header?).
And something is obviously wrong today, since we're still using the same baseline value for our gigabit fiber to the home connections, our 3/4/5G (scratch to taste) mobile phones, etc.
I had to replace my Apple AirPort Extreme when I got gigabit fiber since it didn't have a manual MTU setting and it didn't autodetect the MTU properly over PPPoE... In 2020 I still need to manually set the MTU on my Ubiquiti USG...
Ah, I was always wondering why my ISP configured my fiber modem's mtu to 1492. So it's due to using PPPoE? Is there no way to use bigger mtu when using PPPoE?
Sadly, smaller than 1500 byte MTUs still cause issues for some people to this day. It's all fine if everything is properly configured, or if at least everything sends and receives ICMP, but if something is silently dropping packets, you're in for a bad day. These days, I think it's usually problems with customers sending large packets, as opposed to early days where receiving large packets would routinely fail, but a lot of that is because large sites gave up on sending large packets.
https://www.networkworld.com/article/2224654/mtu-size-issues...
Ethernet uses something similar, but is able to detect if someone else is using the wire, called carrier sense. A short packet of 1500 bytes reduced the likelihood of collisions.
More at:
https://networkengineering.stackexchange.com/questions/2962/...
Another quirky story from the past: Sometime around 20 years ago I was having a bizarre networking problem. I could telnet into a host with no trouble, and the interactive session would be going just fine until I did something that produced a large volume of output (such as 'cat' on a large file). At that point the session would freeze and I would eventually get disconnected. After troubleshooting for a while I identified the problem as one of the Ethernet NICs on the client host. It was a premium NIC (3Com 3C509). Nonetheless, the NIC crystal oscillator frequency had drifted sufficiently that it would lose clock synchronization to the incoming frame if the MTU was larger than about 1000.
US companies had prototypes using 64 bytes, while European companies used 32 bytes. To avoid anyone giving a competitive advantage, they decided on a middle ground of 48.
There were trade-offs between 32 and 64 bytes as well: a 32 byte payload had a higher overhead than a 64 byte payload, but it had a shorter transmission time which made it easier to do voice echo cancellation.
Or so I was told many decades ago when I got introduced to ATM systems...
Even I could of told you, 'the engineers at the time picked 1500 bytes'.
If there are efficiency gains to be had from using jumbo frames, wouldn't setting my MTU to a multiple a 1500 still be of some benefit? If my PC, my switch and my router all support it, that would still be a tiiiny improvement. If the server's network does as well and let's say both of our direct providers, even if none of the exchanges or backbones in between do, that would still be an efficiency gain for ~10% of the link, right?
As a handy feature on Linux at least, you can set your MTU to 9000 locally, and then set the default (internet generally) route to have a MTU of 1500 to prevent issues:
ip route add 0.0.0.0/0 via 10.11.11.1 mtu 1500
Many devices can do over 1500 but anyone who has done so without careful consideration knows the outcome isn't predictable unless everyone on the network is prepared to do so.
A dedicated / controlled SAN type environment can do it just fine, beyond that it can be difficult.
Is my mind playing tricks? Were they faulty units or was there meant to be a crack?
This picture could be the same thing:
https://www.vogonswiki.com/images/3/37/Viglen_Ethergen_PnP_2...
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
I still have a load of them gathering dust somewhere. However a system with ISA on it is a bit rare now and I'm not sure I can be bothered to compile a modern kernel small enough to boot on one. Besides, it will probably need cross compiling on something with some grunt that has heard of the G unit prefix.
More info: https://electronics.stackexchange.com/questions/381682/what-...
Reminds me of the IPv6 adoption problem: https://news.ycombinator.com/item?id=14986324
He’s so optimistic. My brain heard this as “only 20% of packets […] are the maximum size”
What are all of those 64 byte packets? Interactive shells, or some other low bitrate protocol?
The transfer graph is wrong - it shows packet count distribution, not size. Quick math says roughly 90% of transfer size are >= 1024 byte packets.
Wikipedia has this link showing that 9000 bytes was picked by one site c. 2003 simply because it was generally well-supported by their existing hardware: https://noc.net.internet2.edu/i2network/jumbo-frames/rrsum-a...
That is, 9000 is the first multiple of 1500 which can carry an 8192-byte NFS packet (plus headers), while still being small enough that the Ethernet CRC has a good probability to detect errors.
But your PC is still sending more packets.. the modem is struggling to fragment them all and send them upstream.. its memory buffer is filling up.. your computer is retrying packets that it never got a response on..
By lowering your computer MTU to 1492 to start with, you avoid the extra work by the modem, which can have considerable speed increase.
If have a 1500 byte MTU for IP, then we need at least a 1514 byte MTU for IP + Ethernet. We often call the > 1514B MTU the "interface MTU". It's unnecessarily confusing.
But since the local network and the end network where the servers are located will almost certainly be 1500, the point is almost all but moot.
The problem is that Starlink only controls the steps from your router to "the internet". If you're trying to talk to spacex.com it'd be possible, but if you're trying to talk to google.com then now you need Starlink to be peering with ISPs that have jumbo frames, and they need to peer with ISPs with jumbo frames, etc etc and then also google's servers need to support jumbo frames.
Basically, the problem is that Starlink is not actually end to end, if you're trying to reach arbitrary servers on the internet. It just connects you to the rest of the internet, and you're back to where you started.
This is also true for any other ISP, Starlink is not special in this regard.