Why mobile apps suck when you're mobile (TCP over 3G)
blog.davidsingleton.org
blog.davidsingleton.org
Now that smartphone apps are widespread and someone developing a service can control both sides of the connection, there's definitely room for someone to devise a really good TCP replacement (layered on top of UDP) with an iOS library, an Android library, and an Apache mod.
As I recall, one of the key problems this group found was that GPRS had a very low incidence of packet loss. This was because the physical layer protocol had some error correction built in. TCP wasn't designed for this and ended up doing the wrong thing. This group proposed a custom reliable protocol between the mobile device and a proxy gateway, and their results showed improved performance.
We're not talking about GPRS but I am curious why there is so much packet loss?
Also, just FYI: MobiSys (the top academic conference on mobile systems) is happening in the DC area this week.
I have no idea about monetization though. I doubt just charging for the code would work.
The company is now defunct.
One method would be to disable the reliable delivery mechanism in the 3G code.
Another method would be to use huge retransmission timeouts at the TCP level for connections going over 3G, so the TCP level seldom retransmits packets.
If the 3G implementation provides completely reliable connections, the OS could fake the TCP implementation and just use that implementation directly instead of sending TCP packets. (not entirely unlike the sshuttle VPN).
The L1 and MAC of 3GPP are generally implemented in an ASIC and not necessarily visible to the OS. Many of the parameters are dictated and controlled by the RAN (operators are picky and try to control things so that handsets can't misbehave and crap all over the other users in a cell), so the OS doesn't get a look in -- i.e. I don't think it's possible to turn off the 3G error recovery mechanisms.
Disclaimer: I used to work at L1 so someone with more detailed knowledge the MAC and higher layers of the 3GPP stack could give a better idea of what happens when the MAC ARQ limits are hit.
Summary: 3G was designed more than 10 years ago and is showing its age. User requirements have shifted, but standards can't evolve that quickly. Maybe LTE will help provide the user experience for today's apps... question is whether it will be widely deployed in time to matter.
EDIT: updated for clarity and spelling.
If you've got a very simple radio, say, a ham 2 meter handset, you find there are interference patterns caused by multipath interference that cause the signal to get stronger and weaker as you move half a wavelength this way or that way. This causes a fluttering noise when you're listening to somebody transmitting from a car.
A more advanced radio that uses a spread-spectrum signal finds that at any moment of time, some parts of the signal's frequency range is in-phase and some is out-of-phase.
If your goal is to make a low-bandwidth signal robust, you can spread a low-bandwidth signal over a wide spectral range and you won't notice fading. On the other hand, if you're trying to send as much data as you possibly can in the bandwidth you've got, you're ultimately going to implement something like OFDM. Now, in an OFDM system, you're using frequencies in parallel to transmit more data, not to improve robustness. If you're sitting in one place your system can figure out what the fading is and work around it. If you're moving, the relative performance of the subchannels is always changing too fast for OFDM to work.
This is how you can have a good voice connection on your phone (which is using a PHY truly built for mobile operation) but not have it on for data (which is using a PHY built for high performance operation from a fixed point.)
Wireless networks with their higher packet loss (compared to wired ones) fool the sender into believing that the connection is satured. Combine this with the enormous buffer sizes used by mobile broadband systems and the TCP congestion control algorithm throttles the connection to an unusably slow level.
The author is not making a case for trading speed for reliability by using a different PHY, he is making a point for embracing the package loss and delays that exist in current mobile broadband systems and using those characteristics to create a smarter transport layer protocol to replace TCP on mobile devices.
In the big picture, however, it's generally better to fix the part of a system that's broken (the PHY) than it is to try to compensate for the underlying problem in upper layers.
It's also good to be careful about your terms. "Mobile" means operation out of a car or other motorized moving platform. "Portable" means a device that's being carried by a pedestrian. When the bandwidth of a radio channel gets wide, these become very different environments.
All I can say about new protocols is that you can go to your local Uni and find (quite literally) a ton of conference proceedings on this very topic. It seems that somebody has been funding a ludicrous amount of research on TCP replacements for decades without anything practical coming out of it.
Out of all that literature you ought to be able to find something that works or find a good reason why it can't be done.
Let me repeat that: Don't fuck with TCP. TCP is fine. It's BUILT to deal with your lousy network, it's DESIGNED to deal with nuclear attacks from the Guys of the Bad Color, and (ironically) the worst thing you can do to TCP is to try to be nice to it.
"What about keep-alives?"
There are no keep-alives in TCP. If two hosts are in agreement about the address and port, they can keep a connection in an open state until the universe goes cold. (There may be keepalives necessary /above/ TCP, but TCP itself doesn't use 'em). Don't yammer something about your internal state every 30 seconds; you don't do it on the bus home, so don't do it on your network, either.
"Don't we need this proprietary fancy-schmancy buffering algorithm from FuppedUck Communications Consultants? They keep telling us we're toast without it."
No. Just run TCP. Spend money on decent routers.
"I'm feeling really nervous and I want to, like, disable Nagle for everything. Isn't that what you do? Turn Nagle off and maybe muck with the --"
You're from AT&T, aren't you?
It can be fixed with congestion control protocols tacked on the side, I think, but it's designed for wired networking with slow transients in the performance profile.
It IS, however, better than anything that I'd be able to hack together quickly.
Networks are slow. Mobile networks are slower. The most robust fix to the problem is to "optimistically replicate" your application data to the end user's device, so that the network latency does not become part of the user experience.
This is a strong fit for applications like CRM or geographically constrained apps, as the data sets are small enough to fit completely on your devices. For larger data sets the issue becomes: which subset of the data should be copied to the device ahead of time.
The user should never needs to wait on the network. All data operations are played against the local Couch, which handles asynchronously transmitting changes to and from the remote server, in the background. This pattern makes it much easier for app developers to make responsive applications, where users are never left waiting on multi-second round trip times.
Why you'd have big buffers sometimes and not others, I have no idea.
Indian Airtel's network is a live example of that disaster. It is almost unusable, while they still actively promoting 3G and iPhones. ^_^
The product is deployed in some >10Gbps networks, and shows very impressive gains for real live traffic.
Ended up writing a piece on Google because of this on my blog: http://micheljansen.org/blog/entry/1060
(shameless plug :P)
0. http://www.technologyreview.com/communications/21601/?a=f
It's also worth remembering that most of the world has 2G connectivity. Heck, even I'm 2G most of the time (rural UK).
/*
* [...] Note that 120 sec is defined in the protocol as the maximum
* possible RTT. I guess we'll have to use something other than TCP
* to talk to the University of Mars.
* PAWS allows us longer timeouts and large windows, so once implemented
* ftp to mars will work nicely.
*/This is correct. At their furthest apart, Mars and Earth are 22 light minutes apart. Their closest is almost exactly 3 light minutes. For comparison, the Earth and Moon are just over 1 light second apart.
On a more serious note, could tcp.c just be patched on the client to drop packets after ~10sec?
The International Space Station still uses Kermit! At least they did in 2003:
While I don't know that we can fairly say the Internet qua Internet has been extended past Earth (and associated environs), NASA certainly uses a Solar System-scale network already, and while they haven't made a big deal about some of the routing they've already done, if you read the press releases carefully they'll sometimes mention how they routed the signal from one probe through another. It's already a network.
Other trains don't seem nearly as bad. I'm not sure how much of that is due to different construction or (in the case of other modern(ish) rolling stock) anything like the repeaters you mention (though as all the franchise owners are cheap-arses I very much doubt that tech has been paid for by any of them!).
There are a number of places where trains skip around at a goodly speed in what is probably a packed area by way of phone cells, so there will sometimes be a significant number of cell-to-cell hand-offs. Having said that, as I grew up (well, more-or-less) before mobile phones were common I'm still slightly impressed that the whole cell hand-over thing works at all mid-call at 80+mph so maybe they are not much of an issue unless the destination cell is already saturated at the time.
What is really interesting is to watch the RTT while you run a speed test.