Currently working for a bioinformatics data processing company. I've seen, heard and disproved this "many files is better than one big file" argument before... But there is a trick to it...
---
Starting with an example parquet file first, something similar to bioinformatics data i've been working on recently. This is the "fully read into memory as a table / data frame" view of the file
| File ID | spectrum attributes | groupings | numbers and stuff | more numbers as a list | File as raw string |
| -------------- | ------------------------------------------------------ | --------------- | ----------------- | ---------------------- | ------------------ |
| 2984704 | ["other_thing", "another_thing", "different_things"] | group 1 | 329021854.0935902 | [0, 2344, 22, 74, 745] | iw c3lyultrc3l.... |
| 2984705 | ["other_thing", "another_thing", "more_things"] | group 2 | 329021854.0934522 | [231, 09, 123, 15, 5] | sdalkjfh2cn232.... |
| 2984707 | ["other_thing", "other_thing2", "some_thing", "thing"] | group 2 | 3232518.032532 | [892342, 52, 252, 525] | cnm3247cmo27xm.... |
The "File as raw string" column is the magic one here. It's going to let us do what you did, but without having to manually manage thousands, hundreds of thousands or millions of files.
---
### PROCESSING
It sounds like you distributed your computation (smaller pieces of data, executing on many threads). But it reads like you did it manually (it reads like you wrote the orchestration code yourself, rather than sitting there submitting one file at a time). By bunching everything together in one file you can get the tools to do the work for you (at least with current tech you can, no idea about CERN's ROOT).
For modern business, the data engineering tech stack often uses Apache Spark for the distributed data processing engine / cluster engine [0].
Spark distributes data across the cluster nodes by partitioning your data frames/tables/tabular data. Rows 1-500 are loaded on cluster node 1, rows 501-100 on cluster node 2, etc. Spark then executes processing in threads on each node. Each partitioned row on a node is passed to each available thread for that node and processed. The results are stored in memory on the node and can be accessed later on for "other stuff"^{TM}.
Remember that I stored the raw file contents in the "rows" of the "table"? Now I can just "load" that file as part of the processing in a single thread [1]. Et Voila! I'm doing exactly what you did, many files being passed to many threads, but I haven't had to orchestrate anything myself.
Spark has done it for me! So the PROCESSING element can become easier, from the human operator orchestrating the processing of many files perspective, when you have one big file...
What about LOADING the data?
---
### LOADING
> What I think jltsiren is trying to say is that if you're trying to parallelize within one file, you've already lost
I can store this example data as parquet format because it is tabular. As part of the parquet standard, I can:
- partition the data by column values -- a file per partition based on the "grouping" column
- use row groups -- a file per partition of 10,000 row subsets
The above two mechanisms mean you can parallelise the LOAD of the data as well.
When you have a 100TB data set, this becomes *really* important to get right to minimise network data shuffles -- where data is being transferred around the cluster nodes because Spark needs to repartition the data across the cluster. If the data is already partitioned as a parquet file then Spark can load in 1x partition of parquet data onto 1x cluster node.
In an ideal case, where there is 1x cluster node for every 1x data partition, your full dataset will load onto the cluster in the same time it takes to load 1x parquet partition.
Network transfer times for 1TB vs 100TB are significantly different, and this approach can significantly reduce loading time when needing to execute different variants of the processing code on the same source data.
---
In summary, I get where you are coming from. But there are tools that do a bunch of magic things these days so we don't need to worry about stuff.
From jltsiren's original post
> Also, when it comes to parallelization, it's hard to beat running many independent jobs in parallel.
Embarrassingly parallel computation is what everyone is talking about here. Me, you, everyone. Spark does it very well. Python's multiprocessing library does it quite well.
The problem is that no-one thinks about loading and/or storing the data in a convenient format to do embarrassingly parallel computation ... they just stick it in a CSV/TSV file.
---
> This was all done rather straightforwardly with the horrid piece of radioactive software garbage that is CERN's ROOT
https://root.cern/releases/release-62606/
It is still being updated .... I agree with your sentiment, I would not use this out of choice after a brief glance at the docs.
---
[0]: The same principles here apply to something simpler like python's multiprocessing library, which is what I applied to gain an 8x speed up in processing times (they were running it single threaded before).
[1]: See the comment about pymzml about why it's not usually "just" that simple...