Here are a few papers on the problem in chronological order:
1. Original paper briefly talking about the "incast" TCP problem in storage environments: http://portal.acm.org/citation.cfm?id=1049998
2. Our follow up work on that problem from a few years back: Measurement paper ( http://portal.acm.org/citation.cfm?id=1364825 ) and our initial solution ( http://portal.acm.org/citation.cfm?id=1592604 ) with a microsecond retransmission Linux patch here: https://github.com/vrv/linux-microsecondrto
3. Another paper talking about Incast in Datacenter environments, focusing on a different form of the workload: http://portal.acm.org/citation.cfm?id=1592693
4. RAMCloud - a project that briefly talks about the need for low-latency transports: http://www.stanford.edu/~ouster/cgi-bin/papers/ramcloud.pdf
5. ICTCP - paper from last year that tries to solve the Incast problem using receiver advertised window algorithms: http://conferences.sigcomm.org/co-next/2010/CoNEXT_papers/13...
DCTCP tries to go beyond solving the Incast problem and focuses on trying to control buffer occupancy in datacenter environments that contain both long flows and bursty flows.
I think these papers all assume lossy link layers (Ethernet), but there are standards and other technologies (Datacenter Ethernet, Myrinet, Infiniband) that aim for lossless link layers to make the transport problem easier, but come with various other drawbacks today (cost, compatibility, etc.). In the meantime, I hope DCTCP or the microsecond TCP patch prove useful in solving some of these problems.