the problem we were trying to solve included loading thousands of full human genome sequences into an Apache Spark distributed array, which means, behind the scenes, there's a collection of Linux hosts each with a JVM and those JVMs are running Apache Spark code, which shuffles around subsections of the virtual array among Linux (and thus JVM) instances, as the mathematical operators operate on them.
For the genomic data of a single human, we would model our sequencing data as one of (if i remember how i did it) 5 values, for each position in the roughly 3 billion bases we tracked. So the representation of a genome was required roughly 3 billion instances of 3 bits each, but this wasn't how Apache Spark was written to calculate, you end up wasting 2 bits per byte in any semi fast packing scheme, but every byte is being shuttled around endlessly over the network.
We fiddled around with various representation schemes but were always forced into wasting bits that we had no fast means (via processor intrinsics) of moving around.
That's somewhat vague, but hopefully conveys the essence of the frustration.