Van Jacobson Denies Averting Internet Meltdown in 1980s (2012)
wired.com
wired.com
There was much computer vendor hostility to TCP/IP, because it was vendor neutral. DEC had DECnet. IBM had SNA. Telcos had X.25. Networking was seen as an important part of vendor lock-in. Working for a company that was a big buyer of computing, I had the job, for a while, of making it all talk to each other.
Berkeley BSD's TCP was so influential because it was free, not because it was good. It took about five years for it to get good. Pre-Berkeley implementations included 3COM's UNET (we ran that one, after I made heavy modifications), Phil Karn's K9AQ version for amateur radio, Dave Mills Fuzzball implementation for small PDP-11 machines, and Mark Crispin's implementation for DEC 36-bit machines. The first releases of Berkeley's TCP would only talk to Berkeley TCP, and only over Ethernet. Long haul links didn't work, and it didn't interoperate properly with other implementations. (The initial release of 4.3BSD would only talk to some systems during even numbered 4 hour periods because Berkeley botched the sequence number arithmetic. I spent 3 days finding that bug, and it wasn't fun. Casts to (unsigned) had been misused.)
The Berkeley crowd liked dropping packets much more than I did. I used ICMP congestion control messages to tell the sender to slow down, rather than dropping packets. I was more concerned with links with large round-trip times, because we were linking multiple company locations, while Berkeley was, at the time, mostly a local area network user. So they had much lower round trip times.
I'm responsible for the terms "congestion collapse" and devised "fair queuing". I also pointed out the game-theory problems of datagram networks - sending too much is a win for you, but a lose for everybody. This was all in 1984. Today, we have "bufferbloat", which is a localized form of congestion collapse, fair queuing is widely used (but not widely enough), and we have enough core network bandwidth that the congestion is mostly at the edges. Today's hint: if you have something with a huge FIFO buffer feeding a bottleneck, you're doing it wrong. Looking at you, home routers.
Back then, I realized that fair queuing could be turned into what's now called "traffic shaping", but decided not to publish that because it would provide ammunition for the people who wanted to charge for Internet traffic. There were telco people who assumed that something like the Internet would have usage billing. This could easily have gone the other way. Look up "TP4", an alternative to TCP pushed by telcos. That was supported by Microsoft up to Windows 2000.
Berkeley broke the Nagle algorithm by putting in delayed ACKs. Those were a bad idea. The fixed ACK delay is designed for keyboard echo and nothing else. When a packet needs an ACK, Berkeley delayed sending the ACK for a fixed time, in hopes that it could be piggybacked on the returned echoed character packet. The fixed time, usually 500ms, was chosen based on human keyboarding speed. Delaying an ACK is a bet that a reply packet is coming back before the sender wants to send again. This is a lousy bet for anything but classical Telnet. Unfortunately, I didn't hear about this until years after it was too late, having moved to PC software.
UNET was expensive; several thousand dollars per machine. BSD offered a free replacement. So 3COM exited TCP/IP and went off to do "PC LANs", which were a thing in the 1980s.
John Nagle
[1] https://tools.ietf.org/html/rfc896 [2] https://tools.ietf.org/html/rfc970
Nowadays it's a struggle even to get Fragmentation Needed but DF Bit Set ICMPs generated. Too many important routers have to interrupt the main CPU to generate these, and either this then defaults to off or the operators turn it off. Fortunately now we have ICMP-less PMTUD, but it's crazy how long it's taken to get that widely adopted.
TCP/IP was based on datagrams at the bottom levels, and expected them to be dropped if something went wrong, everyone had a phone system mind set - world-wide PTT charged for phone calls, the mindset was that they'd charge for reliable data connections (by the minute or the byte) and the protocols essentially mirrored that - you talked to a supplier who did all the low level stuff for you, and essentially made a call to a remote machine - under IP you sent datagrams and most of them got there.
I think that TCP/IP 'won' (and it wasn't obviously going to win at the time) because:
1) the connection to the wider internet was very simple, and essentially stateless
2) it let people innovate everywhere else up the stack
3) no company (or country) owned it
I didn't see that coming, and I thought things were going to have to be far more efficient.
Especially interesting about TP4 & traffic shaping. The even-numbered hour bug sounds like a nightmare--kind of like cold temperature timer issues I've encountered on micros.
It's personally fascinating to me because in about 1984/5 I was working on a local area network and with a friend 'invented' a connection-oriented protocol that used counters to spot dropped packets (because of Ethernet collisions) and request retransmission. We successfully overloaded the network of about 16 machines using this algorithm as it went crazy retransmitting and upping the collision rate and we began worked on very similar algorithms to fix this (but didn't get that far because had A levels to do).
We were testing tweaks to the algorithm, sending traffic from a university in Germany (where my colleague came from) to our lab in Berkeley. We knew something had gone wrong when we could no longer access the Internet from our lab over the 10Mb/s microwave link due to 90% packet loss. The problem was how to shut down the sender in Germany, and it was 4am there. Even bypassing our microwave problems, we could see that there was also 90% packet loss on the 100Mb/s link out of the university in Germany - the sender was clearly sending around a gigabit - so we couldn't just ssh in. It took us quite a while to find the number of an old-hashioned dial-up modem inside the university and dial in from Berkeley to shut the sender down. Turns out my colleague had inverted the rate function in the code, and when loss increased, it sent faster.
The observations about timers and ack-clocking are still as relevant today.
That's not to say that 1988-style Additive Increase Multiplicative Decrease is perfect as a congestion control scheme. There are lots of issues, ranging from not working well on paths with high bandwidth-delay product, to being overly sensitive to non-congestive packet loss, to causing buffer bloat. But I don't think there's any doubt that the Internet survived growing both in traffic and in hosts by many orders of magnitude, survived all the underlying technologies being replaced multiple times, and survived massive changes in applications, all in part because TCP does a reasonable job of matching offered load to available capacity.