Parallel upload to Amazon S3 with python, boto and multiprocessing
bcbio.wordpress.com
bcbio.wordpress.com
Just use twisted or similar and run as many as you want concurrently with a single thread.
E.g. 'If you have a 10GB text file, it can be faster to split the file across multiple connections when pushing it to S3.' That sounds OK, yeah?
This way takes 10GB read and 10GB write of IO in subprocesses and then running a separate process for every HTTP upload of the individual files.
I'd be surprised if it took more (or even as much) code to do it with a single thread on the original file with as many connections as you wanted using txAWS.
It's the difference between "it works" and "this is a good solution you should use as a model for your own apps."
The problem is just one of language and toolkit abstractions. Too many python APIs are blocking unnecessarily with no way out. A non-blocking version of that API would be obvious how to run in parallel (as the twisted one is).
This is part of the reason for node.js' rise in popularity -- or at least reason for existence. He took the other extreme where nothing blocks. The community is coming up with ways to simulate blocking APIs in efforts similar to what you've done to simulate non-blocking APIs in a world full of blocking APIs.
:)