If you're on a home internet/mobile connection downloading large files, a single download will likely saturate your connection. If you're on an EC2 instance, you should be able to do 10-100x better parallelizing.
Source: I used to work on S3.
More like 80MB/s
My favorite version of this is when you start to use shared memory of some fashion to move terabytes of data from S3 to EC2 to work on it without ever hitting a disk.
Not for everyone, and for sure many times the extra milliseconds saved won't matter, but sometimes you really do need to get hundreds of gigabytes or terabytes of data moved as quickly as possible.
I've been able to saturate 20GB NICs on Ec2 with it (32 cores)
As far as the disk overhead, modern NVMe SSD drives can easily sustain millions of IOPS and multiple gigabytes per second of bandwidth, more than keeping up with a 40 gbps link (such as on a large EC2 instance that does have the connectivity to talk to S3 at that rate).
Edit: I read this recently and if I remember correctly there’s a limit of like a thousand parallel connections to s3
AWS internally probably has higher limits for some of their services, e.g. when you query data in s3 with Athena
There is no reason why multiprocessing for IO in python would use _crazily_ more memory than in an other language, when done properly.
(At work, we had an upload job with ~800k files, ranging from <1kb to >100kb. I looked at rearranging how we stored things to avoid small files, but it ended up a straighter shot to continue to use little files but use a worker pool to make the transfer parallel.)