Pushing the Limits of Amazon S3 Upload Performance
improve.dk
improve.dk
UDP Data Transfer (Open Source): http://udt.sourceforge.net
Bandwidth Challenge: http://www.hpcwire.com/hpcwire/2009-12-08/open_cloud_testbed...
Aspera: http://www.asperasoft.com/
Granted that TCP isn't really "fair" between flows that have different RTTs, one could really justify tuning their TCP behaviour. But the challenge is to do this _automatically_, which is what TCP has so successfully done for the past few decades.
EDIT: DCCP decouples congestion control from reliable delivery and its congestion control algorithms are TCP friendly.
I'm not a protocol hacker, so hopefully one will weigh in here. But I do remember some Sky Is Falling discussion over uTorrent's use of UDP for transfer, and the uTorrent line was always that they were implementing UDP transfer in a way that played nice with TCP.
CLIENT <-UDP-> EC2 Instance <-TCP-> S3
Not certain if it is worth the hassle, though.
Pretty expensive though.
The parallel upload code I'm using is written in Python, using the multiprocessing and boto libraries and is here:
http://github.com/twpayne/s3-parallel-put
It has some nice features, like reading values directly from uncompressed tar files - this means your disk heads will scan linearly rather than seeking around. It can also gzip and set Content-Encoding, restart interrupted transfers from its own log files, and do a MD5 sum check to avoid putting keys that are already set.
Comments and feedback welcome.
OTOH do people ever delete objects from S3? It seems people keep uploading to S3.