Data Center TCP: TCP Congestion Control for Data Centers
tools.ietf.org
tools.ietf.org
Something called DCTCP has been in the Linux kernel since 2014 [1]; that commit even cites what looks like an earlier draft of this RFC that dates back to 2014, although that was on the standards track back then [2]. Why was that effort seemingly abandoned?
[0] https://people.csail.mit.edu/alizadeh/papers/dctcp-sigcomm10...
[1] https://git.kernel.org/linus/e3118e8359bb7c59555aca60c725106...
> This document describes DCTCP as implemented in Microsoft Windows Server 2012 [WINDOWS]. The Linux [LINUX] and FreeBSD [FREEBSD] operating systems have also implemented support for DCTCP in a way that is believed to follow this document. Deployment experiences with DCTCP have been documented in [MORGANSTANLEY].
> Why publish the RFC now, especially since it's not on track to become an Internet standard?
Presumably it's to guide future implementors who wish to attain compatibility with the existing implementations.
One might ask why ever publish RFCs that are not on the Standards Track? But that seems a bit silly. We have Experimental, Information, BCP, and Standards tracks for a reason.
As to why publish now? Most likely because the authors finally got the energy to reach the finish line recently. It does happen that some Internet-Drafts take years to reach the RFC-Editor queue. Taking a long time is not fatal.
https://cloudplatform.googleblog.com/2017/07/TCP-BBR-congest...
BBR is designed for the "hostile internet" where you can't rely on ECN marking, and basically tons of people are willingly and unwillingly plotting against you.. middle boxes that do policing and shaping and just plain bizarre things, routers that clear options, other worse/unfair congestion controls, extreme variation in buffer sizes, etc
The demand for this type of congestion control is seemingly driven by top of rack topology used in cloud architecture.
It would then stand to reason that hardware manufacturers would better solve variable queuing requirements in the switch than developing protocol support known to have a major incompatibility and potentially introduce interoperability problems between vendors in mixed networks.
Second, Linux allows setting the congestion control algorithm per-route. So you could set up DCTCP for communicating with the IPs in the same data center, and use the default CC algorithm for everything else. And what if you can't use/don't want to use per-route settings? Well, you'll generally have two classes of machines anyway. Frontends that can communicate with the outside world, and backends that can't. So you could set up different congestion control based on the role of the machine.
Solving this in the switches seems tricky. Sure, per-flow rather than global or per-port queues could be used to solve the mice vs. elephants problem. But it does not help with TCP incast unless you also add huge buffers. You want switches to be simple, fast and cheap. A switch with per-flow queueing and huge buffers seems like the opposite.
That's one heck of a use-case.
Link: https://www.soe.ucsc.edu/sites/default/files/technical-repor...