Let's make TCP faster
googlecode.blogspot.com
googlecode.blogspot.com
The thing is, a heck of a lot more runs over the Internet/TCP than just HTTP/the web. Also, it can very well be argued that a lot of the "end-user" perceived problems they are trying to fix (e.g. HTTP total request-response round trip latency) are acutally problems with HTTP, rather than TCP - notably the fact that for "small" web requests all HTTP effectively does is re-implement a datagram protocol (albeit with larger packets than UDP) on top of TCP, with all the consequent overhead of setting up and tearing down a TCP connection.
It's an interesting set of fixes. But are they the right fixes, at the right level? Would moving to SPDY instead of HTTP fix the problems better, at a more appropriate level? With less chance of impacting all the other protocols that run (and are yet to run) over TCP?
Changing the fundament of the Internet purely for the sake of a higher level protocol (albeit an important one) in the stack seems dangerous. This would be the case, if for no other reason that it sets a precedent for future changes. Changes at each layer should always be as stack-agnostic as possible. This is by design.
All our work on TCP is open-source and publicly available. We disseminate our innovations through the Linux kernel, IETF standards proposals, and research publications.
(edit: Not so lazy after all I guess. The draft RFC here: http://tools.ietf.org/html/draft-cheng-tcpm-fastopen-00 and after a very quick perusal I don't see an attempt to solve the DOS problem either. It seems like it just requires apps to handle the transactions really fast and then close the connection?)
In my last job creating mobile wireless drivers, we had a problem with wireless roaming. TCP/DHCP are set up assuming IP address establishment is a very infrequent operation. Typically it could take several seconds, which is fine if it only happens at boot or when a human trips over a cable and plugs it back in.
But wireless devices 'plug back in' each time they roam to a new AP. In an industrial environment (warehouse, 60 APs installed over several acres, forklift driving 20MPH) you may need to roam every second or so.
Its time to examine every aspect of TCP for large (huge) installations, very frequent device discovery (power-save in handheld devices), rapidly changing network topologies and so on.
If you select the 'best' as the fastest, then you can simply take the 1st response and run with it. Then it takes only milliseconds as you observed. That's what we did.
http://news.ycombinator.com/item?id=2755461
While it works, and works fast, many people raised concerns about the implications.
A broad summary of the trade-offs are here: http://news.ycombinator.com/item?id=2758576
1. Increase TCP initial congestion window to 10 (IW10).
It seems contradictory with the general concept that too much buffering harms latency and may actually be aggravating congestion: http://queue.acm.org/detail.cfm?id=2071893
Based on our large scale experiments, we are pursuing efforts in the IETF to standardize TCP’s initial congestion window to at least ten segments. Preliminary experiments with even higher initial windows show indications of benefiting latency further while keeping any costs to a modest level. Future work should focus on eliminating the initial congestion window as a manifest constant to scale to even large network speeds and Web page sizes.
https://docs.google.com/gview?url=http://www.cs.helsinki.fi/...
IW10, while improving elapsed times, imposes higher queuing delay than IW3
However, if self-congesting, IW3 is more aggressive in terms of queuing delay
AQM (RED) failed to control the increase in the queuing delayAlso it's much easier as of late to get the benefit from a larger initial cwnd. Back then you needed to recompile the kernel with source tweaks, now you just use a backport or depending on your distro version you already have the benefit as kernel 2.6.39 has the change... http://kernelnewbies.org/Linux_2_6_39
Jim Gettys article "IW10 Considered Harmful" is worth a read too - http://tools.ietf.org/html/draft-gettys-iw10-considered-harm...
TFO basically says the handshake is really just UDP, and the TCP connection doesn't really exist except as a byproduct of an ongoing UDP-based exchange. The 3-way handshake is just the first 3 messages in that chain, and the TCP channel doesn't exist until that many have occured, but the unreliable "phantom" UDP channel doesn't go away once reliability is established. The head/outstanding link in the chain is always unreliable.
I think TCP is a strange mental error: nobody ever needed a to make TCP a real transport protocol next to ICMP and UDP, etc. It didn't need an IP transport number of it's own. TCP is just the idea of "reliability" and can exist entirely in software (and for that reason should, since it's one less thing to maintain in the kernel). UDP is enough. (and ICMP, for example addresses a different problem: out-of-band network feedback.)
Existing code would work the same. I could still ask for a "TCP" connection, and start sending with the real data carried by UDP and benefit from 1 round trip if I don't need to send more.
TFO does that too -- allows some of the unreliability to creep in in the hope that the system is reliable enough that it's worth it -- but it also adds complexity to the existing name "TCP", and I'm not convinced that's good or worth it. TFO solves the right problem in the wrong place IMHO.
There is no additional complexity. This is baked into the kernel.
TCP being in user land in software would be absolutely terrible. There would be many different implementations, it wouldn't be standardised and the fact of the matter is that I as an application developer don't want to have to create TCP on top of UDP. I want to be able to say connect here and establish a connection and make sure my data makes it.
I agree, you as an application developer shouldn't have to recreate TCP. The code already exists, I'm just suggesting that it shouldn't live in the kernel/OS. There's no difference to users or developers at the application layer. (I think evolutionary pressure is a good thing, but there's no reason not to preserve the interfaces for compatibility.)
<?ego_rant("on")?> That said, since we build towers up -- and TCP has already been working for a long time -- it may be against the grain to redirect growth towards the perimeter. It feels retrograde and less snazzy. But if we don't take advantage of the land below us too, the building topples/the goal suffers. Examples would include redundant encapsulation of frames, unnecessary round trips, etc. Start imagining tunneling TCP over TCP (if you've ever forwarded X11 connections over SSH over a 56k modem, you probably know what that would be like). It begins to feel like we're base64-encoding everything.
I think there's an even more important example to think about though.
People jump through major hoops to make their webservers incredibly fast, and able to handle 100s of thousands of connections per second. Worker thread pools, I/O completion ports, you name it. Unfortunately, webservers are serving up TCP connections and TCP needs state to be reliable (otherwise it's just UDP). Unfortunately since TCP is being used to transfer HTTP, which is supposed to be stateless, these goals work against each other.
Imagine how fast a webserver might be if it didn't have to hold onto connection data at all... TFO alone doesn't get you there, it just gets you back to 1 round trip. <?ego_rant("off")?>
I am saying that we wouldn't need to "invent" TFO at this late date if we had started from there (no time like the present). TFO is like digging up though. :)
Also, I would like to see more emphasis given to research on mobile networks, which is my area of interest. Perf for large stable networks is not the same for choppy 3G-ish mobile networks.
After more digging, I was able to find this. If you scroll all the way to the end, there is some verbiage about just setting TCP_RTO_MIN to 1. However, the author claims this causes issues with delayed ACK unless another (missing) patch is applied.
https://github.com/vrv/linux-microsecondrto http://www.pdl.cmu.edu/PDL-FTP/Storage/sigcomm147-vasudevan....
Why does Google? Because web search is behind billions of dollars of revenue. Micro-optimizations matter to them.