DCTCP: TCP optimized for lower latency in Data Centers
stanford.edu
stanford.edu
Here are a few papers on the problem in chronological order:
1. Original paper briefly talking about the "incast" TCP problem in storage environments: http://portal.acm.org/citation.cfm?id=1049998
2. Our follow up work on that problem from a few years back: Measurement paper ( http://portal.acm.org/citation.cfm?id=1364825 ) and our initial solution ( http://portal.acm.org/citation.cfm?id=1592604 ) with a microsecond retransmission Linux patch here: https://github.com/vrv/linux-microsecondrto
3. Another paper talking about Incast in Datacenter environments, focusing on a different form of the workload: http://portal.acm.org/citation.cfm?id=1592693
4. RAMCloud - a project that briefly talks about the need for low-latency transports: http://www.stanford.edu/~ouster/cgi-bin/papers/ramcloud.pdf
5. ICTCP - paper from last year that tries to solve the Incast problem using receiver advertised window algorithms: http://conferences.sigcomm.org/co-next/2010/CoNEXT_papers/13...
DCTCP tries to go beyond solving the Incast problem and focuses on trying to control buffer occupancy in datacenter environments that contain both long flows and bursty flows.
I think these papers all assume lossy link layers (Ethernet), but there are standards and other technologies (Datacenter Ethernet, Myrinet, Infiniband) that aim for lossless link layers to make the transport problem easier, but come with various other drawbacks today (cost, compatibility, etc.). In the meantime, I hope DCTCP or the microsecond TCP patch prove useful in solving some of these problems.
Also, a few catches with DCTCP:
1. It's not completely an end host solution and requires ECN support from switches, which should be widely available. Can someone pitch in about the availability of ECN in their networks?
2. DCTCP and plain TCP don't mix well, and hence not incrementally deployable. DCTCP and TCP+ECN also don't mix well!
I recently talked with someone on the Facebook memcached team about this problem and they mentioned that moving to UDP has been useful, but if I'm not mistaken, I believe they gave up on 100% in-order reliability in return for low-latency.
There have been many improved congestion control algorithms proposed with this train of thought: XCP[1], RCP[2] to name a few. One of the reasons why they have not caught on is that they require a lot of support from the network.
[1] XCP: http://www.isi.edu/isi-xcp/ [2] RCP: http://yuba.stanford.edu/rcp/