Downloading files from S3 with multithreading and Boto3
emasquil.github.io
emasquil.github.io
Technically it is independent from Metaflow, so you could use it as a stand-alone, high-performance S3 client.
See docs here https://docs.metaflow.org/metaflow/data#store-and-load-objec...
And code here https://github.com/Netflix/metaflow/tree/master/metaflow/dat...
(I wrote it originally - AMA if curious)
You might want to update section ‚Caution: Overwriting data in S3‘ in the docs since S3 offers strong read after write consistency since dec 2020.
https://aws.amazon.com/blogs/aws/amazon-s3-update-strong-rea...
Good catch re: the warning about consistency! The docs were written before the change :)
A colleague of mine developed this tool to make this functionality available in a CLI: https://github.com/chanzuckerberg/s3parcp
I've been thinking of starting using Go to deploy some stuff that doesn't need python as dependency, and statically compiled :P
Edit: In this case for DO Spaces. Way more cheap.
RClone: https://rclone.org/ S3 Backend: https://rclone.org/s3/
For instance, if you "cp -a" a directory and then apply sync, it could do nothing and return success if the copied files were last modified before the ones in S3.
For our use case at work, we wanted to be _sure_ that sync always work as intended, and thus ended up recomputing etags locally and compare to the ones in S3 to know what to sync (got bitten by the issue of last modified before)
Just last week I wrote basically the same thing as an ad-hoc solution using boto3 because I had 10s of TB of data to pull out of Glacier and distribute across S3 buckets. It wasn't a big deal because I'm experienced writing parallel network code in Python and having big datastreams flow, and boto3 has good documentation, but things like this really shouldn't be left as an exercise to the SDK consumer.
We have a use case for copying terabytes of content to buckets with different owners and it just seems wasteful to run everything through a client.
I've read over it and I'm reasonably sure that it's going to issue CopyObject, but it would take me actually getting out paper and pen to really track it down.
The AWS CLI and Boto are a case study in overdoing class hierarchies. Not because there's any obvious AbstractSingletonProxyFactoryBean, but rather that there's no instance that stands out as "this is where they went wrong" and nevertheless the end result is a confusing mess of inheritance and objects.
[1]: https://github.com/aws/aws-cli/blob/45b0063b2d0b245b17a57fd9...
I wrote many tests for it... so many tests...
edit: just tried it, definitely seems to be working. not sure what I was seeing earlier, thanks!
At scale, the performance cost is worth it given the number of checks done internally to ensure the copy is perfect. If you download and upload, then you could make it faster however getting all the details right such that corruption didn't happen is tricky.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
You made a great point about going async if you need a working command channel.
I would also point out that for many cases, using a ThreadPoolExecutor works fine.
I maintain a simulation engine and it's distributing work out to AWS Lambda. So I had two distributors, one thread based and one async based. (Also one based on multiprocessing that runs work on a local machine.)
From my tests, the performance is basically equivalent, which is not surprising: most of it was just waiting and then processing incoming responses. Threads work great at this, and botocore is designed to work with threads.
I eventually went with async because Python doesn't let you prioritize threads. That means that if you get a "stop" signal, in the threaded model, the command thread is competing with many worker threads. In the async model, the workers are all in one thread, so the command thread will be woken up per the switch interval[1].
So, broadly, if you're going to need a command channel that must respond in a timely fashion, I'd recommend piling workers into an event loop through async.
The other possibility is to have a command channel run in another process, but then you need to get the fork right, do IPC, etc.
[1]: https://docs.python.org/3/library/sys.html#sys.setswitchinte...
Aws-cli is multithreaded in both download and upload. There is even a setting to tweak the number of parallel requests [2].
[1] https://github.com/aws/aws-cli
[2] https://docs.aws.amazon.com/credref/latest/refdocs/setting-s...
Edit: didn't know about s5cmd (from other commenters). Seems like a faster alternative to aws-cli.
Edit: I read this recently and if I remember correctly there’s a limit of like a thousand parallel connections to s3
AWS internally probably has higher limits for some of their services, e.g. when you query data in s3 with Athena
There is no reason why multiprocessing for IO in python would use _crazily_ more memory than in an other language, when done properly.
My favorite version of this is when you start to use shared memory of some fashion to move terabytes of data from S3 to EC2 to work on it without ever hitting a disk.
Not for everyone, and for sure many times the extra milliseconds saved won't matter, but sometimes you really do need to get hundreds of gigabytes or terabytes of data moved as quickly as possible.
I've been able to saturate 20GB NICs on Ec2 with it (32 cores)
(At work, we had an upload job with ~800k files, ranging from <1kb to >100kb. I looked at rearranging how we stored things to avoid small files, but it ended up a straighter shot to continue to use little files but use a worker pool to make the transfer parallel.)
If you're on a home internet/mobile connection downloading large files, a single download will likely saturate your connection. If you're on an EC2 instance, you should be able to do 10-100x better parallelizing.
Source: I used to work on S3.
More like 80MB/s
As far as the disk overhead, modern NVMe SSD drives can easily sustain millions of IOPS and multiple gigabytes per second of bandwidth, more than keeping up with a 40 gbps link (such as on a large EC2 instance that does have the connectivity to talk to S3 at that rate).
It was probably a good article tho
RClone: https://rclone.org/ S3 Backend: https://rclone.org/s3/