- Arrow is an efficient object format that helps skip serialization
- Spark works in memory
- Minio makes storage of data scalable and cheaper than hard disks
- (from the article) Spark jobs don't (didn't??) saturate the network
- (from the article) Spark jobs needed to write the same data to many buffers to get anywhere (exporting?)
- memory mapped IO is here to help (this seems like they're mixing memory mapped device IO
Is the gist that because arrow does not require much serialization/deserialization, memory mapped network devices can shoot it out to another server (w/ S3 as an intermediary) at network speed, skip the disk, and get used at the destination faster? I guess the other method was program-1 -> disk -> network -> disk -> program-2 and now it's program-1 -> s3 -> program-2 ?