TCP's congestion control saved the internet
theregister.com
theregister.com
Still, it has its limits – the most obvious one to me is mentioned in the article: The optimal buffer size for each hop that TCP (or similarly flow-controlled traffic) traverses depends on the average delay and bandwidth, but these are per-flow quantities which an intermediate router can't properly estimate in principle.
I've played around with Linux's excellent fq_codel a bit with promising results in the past, but ultimately, that's not a solution in the spirit of the Internet (dumb routers, smart endpoints) since it depends on looking at individual flows, something routers aren't really supposed to have to do.
I really hope that something like Google's BBR will become the de-facto TCP congestion control algorithm in the end.
Me too. We really don't know what to do about congestion in the middle of a pure datagram network. What saved the Internet was cheap fiber backbones. If backbone bandwidth had remained expensive, this would never have worked.
When I was working on congestion in the early 1980s, our goal was to run at about 70% long haul link utilization. We would back off quite a bit if things got anywhere near saturation. This was fine with the DoD people who funded us. They wanted reliability under stress, not maximum performance at minimum cost. Much of the work on congestion since then has involved trying to squeeze out that last 30%, at the cost of much-increased complexity.
The "bufferbloat" mess is embarrassing. If you use FIFO queuing into a bandwidth choke point, latency is going to go way up. This is well known. It is still ignored in most cheap routers. Every home router feeding limited bandwidth should have outbound fair queuing. Looking at individual flows near the network edges is a big win.
the fq part does most of the work and is intended for edge-routers as bufferspace requirements diminish with growing number of flows.
We have now produced variants of the flow queuing algorithm for fq-codel, fq-pie, and the FQ in cake, which we are now using in libreqos. It is the best form of fq ever designed, IMHO, with really good min/max fairness properties, resistence to starvation, great rtt fairness and highly complementary simultaneously time based AQM technology easily applied. I would like to see more improvements on the AQM side in the future and for it to land in more hardware.
Admittedly, I´m one of the authors... :)
flow-queuing for the mainstream is a godsend in the information age.
but: Eric Dumazet came up with the flow queuing fast/slow queue concept in a single saturday afternoon in his hammock (after admittedly 9 months of hard work on every other fq and aqm algorithm), slammed it into the kernel the next day, and has been building ever more amazing things ever since.
I try to make sure people understand that I am more the acolyte of kathie, van, and Eric, than inventor here. I was proud of this thing we´d developed called SFQred, and writing a paper on it... but then kathie and van came out with codel a month later... obsoleting that part...
Eric and I implemented it over a week, saw that it was far better than RED... I went to sleep for a week, only to wake up to fq_codel, which solved everything.
We have been just re-implementing it and improving the edge cases since. I am pretty happy with cake, now. The world evolves however, and there are new demands on queuing. I cover some of these here: https://docs.google.com/document/d/1tTYBPeaRdCO9AGTGQCpoiuLO...
2. fq_* is meant for end-/edge-routers where the number of flows is rather small
https://web.stanford.edu/~sarslan/files/Updating_the_Theory_...
i assume performance optimizations for the fast-path were successful, or is this scale-up rooted in advances with multithreading and packetsteering?
src here: https://github.com/LibreQoE/LibreQoS/
You can get a transparent high speed bridge with shaping capability up pretty rapidly with supported hardware.
I still long for an ethernet card that can do a trie lookup natively! The flamegraphs are mostly getting the right packet to the right cpu, still.
FQ solves most of the problem thoroughly enabling delay based transports, Van and Kathie´s codel, most of the rest. Van was also on the team that produced BBR, which spends most of it´s time in delay based mode when fq_codel is on the path.
I do wish more folk were actively applying cake nowadays. It works wonders.
https://blog.cerowrt.org/post/juniper/
Libreqos is doing nice things with cake.
It's hilarious to see the things which were hard coded in the first BSD implementation, like a 4 KiB socket buffer size.
Also, as far as I can see, Bill Joy DID actually have linear backoff but it was hidden behind an always-false condition, so his hand-written lesser float backoffs were chosen instead. This seems like a fairly tragic coding error to me.
I have code like that. A return then some body of sane looking but outcome unknown code.
I use git to sync work between multiple computers, and I want to be able to sync work-in-progress code that should absolutely never see the light of day, so I feel like committing and pushing it to a WIP branch is fine.
Even if you only use one computer, I think it's a good idea to push these things so that if your HDD fails you can pick back up where you left off, even if it's code that's not ready to actually merge yet.
It's still used in the telco space; xDSL runs over ATM, generally speaking.
It's not obvious that TCP/IP would win. I wonder how much of its victory was because it was basically free; the specs were public, and pretty much everyone build an implementation.
Compare that to the other protocols, which tended to be stuck to a vendor (DECnet, AppleTalk, IPX, token ring). That makes interoperability difficult, since in general companies didn't want to license their tech to someone else because it was a competitive advantage.
You can still find the free TCP/IP implementations out there (like tinytcp, which is still floating around).
[0] https://www.cs.columbia.edu/sip/articles/tdc0598cover1_side1...
However the main long-staying effect of MPLS is enabling L2 and L3 VPNs (VRFs).
Ethernet hubs were much cheaper than anything else. This was the big win. At that point, the only question was TCP/IP or IPX/SPX (Novell).
You could use Ethernet to hook a couple computers in a room together directly and it worked (mostly). Every other technology required some expensive thing to sit in the middle and direct it all.
Then, when you needed to connect a couple rooms together, you could throw a cable over a wall and hook them together.
Finally, when you got tired of debugging the inevitable continuous failures due to the "idiots in that one room", you threw an Ethernet hub into the mix to hook the rooms together. This was a bit more expensive, but still not as expensive as the idiotic middlebox for the other technologies. And the improvement in reliability and debugging was palpable enough that your boss didn't give you too much static. This was THE big win that other technologies couldn't deal with.
Finally, when you had enough computers that even hubs just weren't cutting it, you bought an Ethernet switch. Yeah, that Ethernet switch was about as expensive as the middlebox from all the other technolgies. However, by now, Moore's Law had been running for a few years so an Ethernet switch was simply expensive rather than outrageous. And you had all this installed Ethernet kit, so your sure as hell weren't going to rip it all out.
At this point, the installed base made sure that Ethernet was going to win.
ATM-based DSL was a compromise of the different factions inside telcos. The 'net heads' got a fast Internet access technology. The 'Bell heads' got ATM into every home in preparation of all the glorious ATM applications that would be available any moment now.
CSMA/CD works pretty well on a shared wire, but not very well on a single radio frequency where you cannot "hear" other people trying to reach the same mountain-top receiver. I'm glad that aspect of congestion is now a thing of the past. (Now that everything (wired) is switched, it's all point-to-point, so the only congestion we get is from insufficient bandwidth on the trunk, and we have ECN for that.)
Yes. As I wrote in 1985, in [1], under "Game Theoretic Aspects of Network Congestion" and "Fairness in Packet Switching Systems", about fair queuing,
We would like to protect the network from hosts that are not well-behaved. More specifically, we would like, in the presence of both well-behaved and badly-behaved hosts, to insure that well-behaved hosts receive better service than badly-behaved hosts. We have devised a means of achieving this.
The goal of fair queuing is not to improve network performance overall. The goal of fair queuing is to reward well-behaved hosts over badly-behaved hosts. If everyone is well-behaved, the queue lengths are the same, usually 1 or 0, and fair queuing does little.
There is an inherent conflict between this goal and achieving maximum data transfer rates. If you try for near 100% utilization, the problems become much worse. You can run comfortably at maybe 70%. This was an accepted tradeoff for DoD systems. DoD wants things to keep working in a crisis, even if normal operation is a bit slower.
This is why I'm not a big fan of HTTP/3. It's a attempt to get about 10% more performance in the good case, at the cost of considerable extra complexity and less immunity to gaming the system.
I never wrote about that much at the time, because if I had, people would have realized earlier that traffic shaping is possible, which implies that you can sell and bill for bandwidth and quality of service. We might have ended up with pay per packet.
To be clear, are you mainly complaining about the multiple streams aspect here, or do you have other concerns? Because while I'm no expert, using HTTP/3 (or QUIC, I'm ignoring the terminology differences here to avoid confusion with pre-standard QUIC, though I guess maybe we're past that point by now?) with just a single stream seems to be the only scalable solution to a certain problem.
The main problem I've had with TCP is that anybody can maliciously inject an RST, and for long-lived active connections, the chances of this happening approach 1. If you control both ends you can just tell your firewall to drop all RST packets so the applications can keep talking to each other, but ...
Dealing with this at the application layer is ... marginally possible I guess, but not easy unless your application looks a lot like a web browser (and even those tend to usually force the end user to manually fix it). Creating a new TCP connection sucks and you have no idea how much data got dropped after it left the application's buffer.
So we need a UDP reliability layer, since that's the only other general-purpose protocol that actually gets routed (and SCTP is flawed anyway). And dozens of those exist ... but almost none have major use, so I can't trust that they will play nice with the wider internet.
With HTTP/3, if there are major flaws found, I have confidence that - like TCP - it has enough users that someone will come up with an acceptable solution. (admittedly, since it's not part of the kernel it will be harder to actually deploy the fix)
(I guess TCP-over-VPN is also viable for some users)
https://github.com/muxamilian/fair-queuing-aware-congestion-...
Mitigations include high resolution retransmit timers (RTO) and FQCN.
Didn't help: SACK, TCP Reno, and New Reno.
[0] https://en.wikipedia.org/wiki/TCP_global_synchronization
https://systemsapproach.substack.com/p/how-congestion-contro...
What would the internet have looked like without Van?