Advanced Scientific Data Format
github.com
github.com
Simply uncompressing a file can be a significant bottleneck, as this can inherently use only 1 CPU core.
Keep in mind that these files are intended to be processed on monstrously huge 128 core and 2TB memory machines! What happens is that the system uses 0.5% of its capacity to manipulate the files while the other 99.5% is heating the data center air.
I'm looking at the CPU usage metric graph now, and the machine is spending half the time at 100% load and half the time at 1% load. That's half the total capacity wasted.
If you find yourself in 2022 or later designing a file format intended for bulk data and you use any of the words "stream", "serialization", or "text", stop. Rethink what you have done, and this time consider that normal machines you can buy for normal money have 256 hardware threads. Soon, 512, and then 1024 in just a few years!
Everything at this scale should be split into blocks for parallel processing. Shoving text files into a .tar.gz archive is not acceptable any more. I don't care what 1970s "standard" it adheres to, that just doesn't matter any more, the hardware has moved on.
I think it's high time that the industry standardised on a generic "container" format to replace legacy archive file formats. Something akin to a random-access archive file like zip, but designed so that even a single file can be decoded in parallel, even if compressed.
As an example of "parallel thinking", CRC or SHA style checksums are automatically a no-go. They're inherently sequential. Instead, Merkle trees would have to be used.
Compression efficiency of small files with random access could be improved by using something like zstd as the internal compression algorithm, but with a shared dictionary stored separately. This would retain the advantages of 'tar' without requiring sequential decoding of files.
Etc...
One key problem is that technology changes quickly. There are always new instruments generating new kinds of data with new properties and new features. People are using that data in new applications.
Software comes many years behind the state of the art. First you need to figure out what is the exact problem the software is supposed to solve. Then you have to solve the problem and turn the prototype into a useful tool. This work is mostly done by researchers who may know a little about software engineering. By the time the situation is stable enough that software engineers who are not active researchers in the field could be useful, it's often too late to change the file formats. There is already too much legacy data and too many tools supporting the established formats.
Another key problem is that the "broken" file formats are often good enough. When you have tabular data where the fields can be reasonably understood as text, a simple TSV-based format often gets the job done. Especially if individual datasets are only tens of gigabytes. By using a custom format, you avoid having to choose from many existing formats that all have their own issues. And that often guarantee you version conflicts and breaking changes in the future.
Also, when it comes to parallelization, it's hard to beat running many independent jobs in parallel. While computers are getting bigger, individual problems are often not, as the underlying biological problems remain the same. In the work I do, a reasonable target system had 32 cores and 256 GB memory in 2015. That's still a reasonable target in 2022. The computers I use have become cheaper and faster, but they have not really changed.
While using Apache Spark for bioinformatics [0] never really took off, I still think Parquet formats for bioinformatics [1] is a good idea, especially with DuckDB, Apache Arrow, etc. supporting Parquet out of the box.
Those upstream tasks tend to be row-oriented. You often iterate over all rows, do something with them, and output new rows in another format. Alternatively, you read the entire input into in-memory data structures, do something, and later serialize the data structures. Using column-oriented formats for such tasks does not feel natural.
Each of your plateaus of stability often need to become recognized before the next step can be taken.
With scientific software at the end of the train, the file-type/file-system needs to be well established and more stable than any software could be, and for a lot longer than the whole scientific project itself takes.
A scientific filetype needs to be well-documented (better than ordinary software) and unchanged for long enough so that all agree no further changes are intended. It needs to have already been virtually perfectly stable, for more years than most research projects are likely to have their data remain useful in the future.
netcdf is in this category while still being extensible and it is very old (well established) if not well known. Public domain government "codec" which basically decompresses netcdf to structured text, and in reverse.
Now this new ASDF filetype looks like it does have useful features of its own except one thing:
>ASDF is under active development
Which can still be a drawback in this situation.
In what other industry is less than 0.5% utilisation accepted as "gets the job done?"
TSV is a terrible format for multi-gigabyte files, because it uses line breaks.
Technically it's possible to parse them in parallel, but if the format enables quoted strings with newlines, then it isn't possible to do this safely.
A simple binary row-based or columnar format could be read orders of magnitude faster. No character-by-character processing needed, just "memcpy" and go.
This matches my experience a decade ago as a rare scientist who could program. When I had to run a job on the giant cluster against our ~100TB dataset, I did not put hundreds of threads to work against one file. We had everything broken up into hundreds of files, so I could run one thread against each file (with a tiny bit of boundary patch-up), and it all just flew. Trying to get it "all in one" would have been an exercise in misery and pain.
This can also be done by simply running many jobs at once. Many (but not all! never all!) scientific analyses naturally work well with that model, so don't fight it on technical purity grounds.
This was all done rather straightforwardly with the horrid piece of radioactive software garbage that is (was? please say was? I dare not check) CERN's ROOT. It had few redeeming characteristics... except for being fast and efficient on massive datasets, once you got it running at all. And that counts for something!
I disagree, you don't need to split up a file to parallelize things if you just use a moderately recent format.
Put it in parquet, have sensible row groups, and turn on zstd compression. Split it into multiple files if you want but fast access to subsets of files which are neatly compressed is very easy to get now.
You also get the ability to store things losslessly, which you don't get so much without custom work with TSV (main example here is floats) and things like schemas.
I've just tested out moving from one of my zstd parquet test files that's 190M, and it's 4.8G as a CSV file for a single datapoint.
Yes, not everyone has such large data sets. (And not everyone has such small data sets!) But it is imperative that scientific computing infrastructure developers understand what classes of experiments they are serving, and what classes they can not or should not serve.
But yes, it is better now than back in the 4.X days or whatever ungodly version it sounds like you used.
---
Starting with an example parquet file first, something similar to bioinformatics data i've been working on recently. This is the "fully read into memory as a table / data frame" view of the file
| File ID | spectrum attributes | groupings | numbers and stuff | more numbers as a list | File as raw string |
| -------------- | ------------------------------------------------------ | --------------- | ----------------- | ---------------------- | ------------------ |
| 2984704 | ["other_thing", "another_thing", "different_things"] | group 1 | 329021854.0935902 | [0, 2344, 22, 74, 745] | iw c3lyultrc3l.... |
| 2984705 | ["other_thing", "another_thing", "more_things"] | group 2 | 329021854.0934522 | [231, 09, 123, 15, 5] | sdalkjfh2cn232.... |
| 2984707 | ["other_thing", "other_thing2", "some_thing", "thing"] | group 2 | 3232518.032532 | [892342, 52, 252, 525] | cnm3247cmo27xm.... |
The "File as raw string" column is the magic one here. It's going to let us do what you did, but without having to manually manage thousands, hundreds of thousands or millions of files.---
### PROCESSING
It sounds like you distributed your computation (smaller pieces of data, executing on many threads). But it reads like you did it manually (it reads like you wrote the orchestration code yourself, rather than sitting there submitting one file at a time). By bunching everything together in one file you can get the tools to do the work for you (at least with current tech you can, no idea about CERN's ROOT).
For modern business, the data engineering tech stack often uses Apache Spark for the distributed data processing engine / cluster engine [0].
Spark distributes data across the cluster nodes by partitioning your data frames/tables/tabular data. Rows 1-500 are loaded on cluster node 1, rows 501-100 on cluster node 2, etc. Spark then executes processing in threads on each node. Each partitioned row on a node is passed to each available thread for that node and processed. The results are stored in memory on the node and can be accessed later on for "other stuff"^{TM}.
Remember that I stored the raw file contents in the "rows" of the "table"? Now I can just "load" that file as part of the processing in a single thread [1]. Et Voila! I'm doing exactly what you did, many files being passed to many threads, but I haven't had to orchestrate anything myself.
Spark has done it for me! So the PROCESSING element can become easier, from the human operator orchestrating the processing of many files perspective, when you have one big file...
What about LOADING the data?
---
### LOADING
> What I think jltsiren is trying to say is that if you're trying to parallelize within one file, you've already lost
I can store this example data as parquet format because it is tabular. As part of the parquet standard, I can:
- partition the data by column values -- a file per partition based on the "grouping" column
- use row groups -- a file per partition of 10,000 row subsets
The above two mechanisms mean you can parallelise the LOAD of the data as well.
When you have a 100TB data set, this becomes *really* important to get right to minimise network data shuffles -- where data is being transferred around the cluster nodes because Spark needs to repartition the data across the cluster. If the data is already partitioned as a parquet file then Spark can load in 1x partition of parquet data onto 1x cluster node.
In an ideal case, where there is 1x cluster node for every 1x data partition, your full dataset will load onto the cluster in the same time it takes to load 1x parquet partition.
Network transfer times for 1TB vs 100TB are significantly different, and this approach can significantly reduce loading time when needing to execute different variants of the processing code on the same source data.
---
In summary, I get where you are coming from. But there are tools that do a bunch of magic things these days so we don't need to worry about stuff.
From jltsiren's original post
> Also, when it comes to parallelization, it's hard to beat running many independent jobs in parallel.
Embarrassingly parallel computation is what everyone is talking about here. Me, you, everyone. Spark does it very well. Python's multiprocessing library does it quite well.
The problem is that no-one thinks about loading and/or storing the data in a convenient format to do embarrassingly parallel computation ... they just stick it in a CSV/TSV file.
---
> This was all done rather straightforwardly with the horrid piece of radioactive software garbage that is CERN's ROOT
https://root.cern/releases/release-62606/
It is still being updated .... I agree with your sentiment, I would not use this out of choice after a brief glance at the docs.
---
[0]: The same principles here apply to something simpler like python's multiprocessing library, which is what I applied to gain an 8x speed up in processing times (they were running it single threaded before).
[1]: See the comment about pymzml about why it's not usually "just" that simple...
Independent jobs go beyond what is usually understood as embarrassingly parallel. In a typical bioinformatics workflow, you download the data, process it locally for hours, and send the results back. The ratio of computation to data is high enough that data transfers (and reading/writing data) rarely become a bottleneck. Meanwhile, the natural size of problems is both large enough that handling them takes hours and small enough that they can be handled on commodity servers.
In work like that, distributed systems tend to be overengineered solutions that make everything harder and more expensive. They make it harder to find developers who understand both the technology and the business needs – the biological problems in question. And they make it harder to install the solution in a new environment that may be fundamentally different from the one it was developed in. Especially in the fairly typical case where the person trying to install it does not have administrator rights.
Independent = embarrassingly parallel, i.e. CPU bound. You want to apply some f on x to get y ==> y = f(x). You have many cases of x, all with f applied independently to get all the independent y outputs. There is no difference between "independent" and "embarrassingly parallel" in this case.
> In a typical bioinformatics workflow, you download the data, process it locally for hours, and send the results back.
And I'm saying that you do not need to be waiting for hours, if you do the data storage and data processing using sane and modern methods (i.e. not with CERN's ROOT tool).
This is not theory BTW -- I've been doing it for the last month. This is a very real, evidence based observation.
---
I was writing out a bunch of other point by point replies here, but I get the sense that it's more likely going to cause you to dig in to your currently held beliefs rather than open you up to the magic methods used by data engineering teams, so I decided to give up and go eat some pizza.
I enjoy pizza.
We are talking about the kind of bioinformatics tools that establish new popular file formats. They are typically developed by researchers rather than software engineers. These tools are intended for end users to install and run on a wide variety of systems. And they are often released before the first high-profile papers on the topic are published. At that point, it's rare to find a software engineer who is familiar enough with the topic to be able to contribute.
Waiting a few hours for the results is not a big deal with these tools, because you are probably going to submit a large number of jobs anyway. Getting the full results will likely take days, regardless of whether you are using naive or modern methods. Because you are probably not allowed to use unlimited compute, your primary constraint is throughput rather than latency. Naive file formats survive, because tools that solve independent problems locally provide similar throughput to state-of-the-art distributed systems.
Bioinformatics, and particularly the subfield that focuses on genomics, is noteworthy because it relies more on freely available open source software than most other academic fields. That may be because the state of the art is changing so rapidly.
Companies sometimes take over old well-established problems. By providing services instead of tools, they can use whatever technology they prefer internally. But you rarely see these companies attacking state-of-the-art problems.
If you have a 50-gigabyte TSV file, you can memory-map, index, and validate it in a few minutes. Overhead like that is usually negligible. And if you encounter something weird like quoted strings, you can just print an error message and abort. It's your file format, so you don't have to make your life difficult by supporting unnecessary features.
That approach still kind of works with hundreds of gigabytes of data. Once you have terabytes, text files become too cumbersome. (Here I'm unfortunately speaking from experience.)
Put it into a database if you care about these things. But let's not reinvent the wheel, and let's not bikeshed binary file formats, of all things. It's a non-starter. Plain text is perfectly fine for long-term storage and interop, converting to a more effective representation is just a one-time cost.
As an example, consider clinical use cases. It can be near impossible to update the software in the pipelines for these situations. How do they move on from the old tools? Ok, now suppose they need to interact with data from non-clinical sources?
Things like you're proposing can be done. As others have intimated in this thread, solving the technical issue is the easiest part of the chain. More than one Tech Person has walked into the fray assuming the only thing holding stuff back is that no one ever thought to apply Good Tech Solutions. Instead what's necessary are people who can understand why things are the way they are, and the human dynamics at play. Building a better mouse trap requires taking these into account from the start.
Interesting discussion though. I write some academic code in which I anticipate file load will be a bottleneck when we scale up one day, so I appreciate your thoughts
It's not that hard to pre process prior to feeding into whatever the latest processing architecture of the decade is .. but it's hell week when you're tasked with decoding some decades old non standard "only used for 18 months" "best practice at the time" obscure compression format data.
Field seperated line orientated ASCII data doesn't require extensive out of band notes to understand or decipher years after you've passed on.
That's quite disingenuous, there are very well documented data formats that have been around for decades (netCDF, HDF) those are not "obscure" formats and are infinitesimally better than text.
> Field seperated line orientated ASCII data doesn't require extensive out of band notes to understand or decipher years after you've passed on.
Field separated ACII data is terrible because it does not contain relevant metadata (except for column names), you quite possibly loose precision (also what even was the precision?), are slow to read, need to be additionally compressed ...
so, not much better?
Sounds like they better stick with the TSV format then ;)
* text/TSV data files made public as a requirement from the published articles may have spaces or dots in column names, missing line ends, non-unique row IDs, etc.
* there is no limit as how many records can be stored in a single file. Latest dbSNP has more than 1000 millions of rows.
* bunch of formats (GFF, GTF, VCF) has several TSV delimited "proper" columns (as: 1 value) and then special column where optional fields are piled in a different format, with another separator, field names etc. Real fun to parse...
It's very optimized towards my specific needs but could be a basis for what you mention
AFAIK the linearity of CRCs allows them to be split, parallelized, and combined. This is also the case for polynomial MACs, like GHash or Poly1305.
Besides, "bioinformatic formats" is a meaningless word anyway. FASTQ, VCFs, BCL, AIRR-seq -- all different and it just works.
Also, I'd love to see someone open a 75 GB FASTQ file in Excel.
FASTA (and its various incantations) are not going anywhere anytime soon.
(I actually know biologists who have run into this problem.)
Abeysooriya, Mandhri, Megan Soria, Mary Sravya Kasu, and Mark Ziemann. “Gene Name Errors: Lessons Not Learned.” PLoS Computational Biology 17, no. 7 (July 30, 2021): e1008984. https://doi.org/10.1371/journal.pcbi.1008984.
Every standard nowadays aside from the very first one, are an N+1. Heck, even IFF and ASN.1, the absolute old timers of file/serialization formats, are improvements on "just mmap to disk" application formats.
And this is the crux of the issue, people still think excel processing is acceptable practice in 2022. If you are required to publish your analysis code (if you are not yet, it will come, the writing is on the wall), are you just publishing the excel sheets?
Look this up: gnu parallel
Also, with building on what the top reply to what you said, if "uncompressing is a bottleneck" and you didn't know you could uncompress multiple files at once suggests you should spend time learning about tools that exist before you jump and try to force your colleagues (who likely have much more experience than you) to adopt "new practices" that the js community will move on from in 6 months.
The resulting CSV file can't be processed in parallel, as there might be quoted record delimiters.
There are compressed file formats where a single file can be processed in parallel. Apache Avro and Parquet are examples I'm familiar with, but these handle columnar data (i.e. replacing CSV), not large numeric matrices etc.
The OP says they get to "ten gigabytes in size." It does not take 15 minutes to decompress a 10GB file on a modern workstation, as they complain in a child comment, unless you decompress multiple files sequentially. I've routinely compressed and decompressed 100GB+ files on 5 year old workstation class machines and it takes at most 5-ish minutes for one direction, (stress on "at most").
Nit: I think I know what you are getting at, with blocks streams vs byte streams, but it's kinda hard to design a file format without serialization or byte streams. Not sure how that would work.
> I think it's high time that the industry standardised on a generic "container" format to replace legacy archive file formats.
I have a side project chipping away at just such a thing. It's quite daunting, so if this at all interests anyone, please comment/reach out. I'd love more of an excuse to work on this.
SITO in a nutshell:
- It's all based on msgpack, which does most of the heavy lift for serialization and datatype encoding
- a sito stream comprises blocks, each block is a self-contained, independently decodeable msgpack array object.
- each block an array of the form (type: smallint, header: optional(hashmap), data: any)
- the type is either a single-byte int, or a packed int indicating a substream id
- the sito primary stream comprises multiple independent substreams
- there's no raw plaintext fields, but there is a plain unicode block which can be used to embed whatever plaintext metadata
- substreams can each have whatever compression/codec they want
- since each stream is a block stream, it's trivial to de/interleave
- for data integrity, I want to do something like block-level CRC/FEC along with per-stream merkle trees but I haven't worked out the details yet
- there are periodic "sync-blocks" which have a magic 8-byte sequence for starting a file, but also throughout the stream, to facilitate re-alignment of read heads
- I've also been toying with the idea of using sqlite as stream indexes and as a general glue to keep track of what's going on (right now, you can arbitrarily start a new substream at any point in the primary stream, so it's hard to tell at the start of a file what's in it, sqlite pre-allocates pages so write heads can go back and update a prior index block)
- nd-arrays are a particularly interesting datatype so there's an emphasis on ergonomics around handling them
I plan on doing a simple PoC at some point soon showcasing SITAR, the sito archive format, with a python tarfile-like interface.
I recall going through ASDF, BSDF and a handful of other formats, finally ending up with HDFS -- which was okay, but not fully satisfactory (I don't recall all my gripes right now).
FASTQ is a plain text format, and is usually gzipped. This is usually not a problem, as the only thing that's going to happen to a FASTQ file is you're going to shove it through an aligner, and a single thread can un-gzip that file fast enough to keep a lot of cores busy doing the alignment.
SAM is basically used for nothing, except very briefly as an output from the aligner before it is promptly converted into BAM.
The BAM format is actually sensible. It's compressed and indexed. There's an alternative format out there called CRAM, which can be a little more efficient, but you need to ensure that some external files are still available in order to decompress it.
VCF (and gVCF) files are text, and they are usually gzipped. Whether they are gzipped or not, they usually have an accompanying index file that allows any section to be accessed without having to sequentially read through. This is possible with the gzipped version because a variant of gzip called bgzip is used that compresses blocks of data, and locations of the start positions of those blocks are stored in the index.
The DepthOfCoverage file format - OK, I don't have any defence of it. It's just huge. If you're storing the read depth for a single sample, it uses about 24 bytes per location, to store a single number that's usually less than 256. So, a typical file for a whole genome sequencing sample would be around 75GB. It also has no index. A few months ago I decided to write an alternative file format, which delta-encodes then huffman-encodes the number in blocks, and uses about 1.3 bits per location, and includes an index, so that 75GB file is now 0.5GB and is a heck of a lot faster to read.
As an aside the GATK DepthOfCoverage is a fairly dire example of slow software. About 7 years ago I wrote my own DepthOfCoverage, which produces the same results but runs about 50 times faster. It wasn't hard. And because the BAM file format is sensible, yes it does access the files in parallel as you suggest.
There are advantages of text file formats. The format is unlikely to be non-readable in 10 years. You can just load it up in less and have a read. The text format doesn't stop it being compressed and indexed and accessed in parallel. But yes, the data could often be stored in a more efficient manner.
Finally, if you're finding that your server is regularly under-utilised, then you aren't doing load-management properly. You should use a queuing system that knows for each job how much RAM and how many CPU threads are used, and therefore how many can be run simultaneously.
CRC is 100% parallelizable FYI. Both in SIMD and also on a block level which can be merged.
import asdf
import numpy as np
x = np.array([[1, 2], [3, 4]], order="C")
y = np.array([[1, 2], [3, 4]], order="F")
tree = {"x": x, "y": y}
af = asdf.AsdfFile(tree)
af.write_to("example.asdf")
and you get in the metadata no distinction between the two arrays even though things like byteorder are included:
x: !core/ndarray-1.0.0
source: 2
datatype: int64
byteorder: little
shape: [2, 2]
y: !core/ndarray-1.0.0 source: 0
datatype: int64
byteorder: little
shape: [2, 2]
This makes me wonder what it's actually storing - is it actually doing something like pickling the NumPy array?You can't infer the stride from the raw binary; a 2-D array [[1, 2], [3, 4]] in C ordering just looks like:
1,2,3,4
and in Fortran ordering it's
1,3,2,4
So there must be additional metadata stored. Maybe it's just using the .npy format internally - https://numpy.org/devdocs/reference/generated/numpy.lib.form...
Frankly, I feel like making the metadata of a scientific data structure human-editable is something of a mis-feature, or at best a non-feature. I use metadata in HDF5 files as a form of provenance tracking and I'd rather there be some friction to editing it.
[0] https://asdf-standard.readthedocs.io/en/1.0.3/intro.html
I would guess that HDF5 would be the better choice for large datasets. However I quite do not understand the capital 'Tree' in this sentence and what that means for practical data sets.
https://www.sciencedirect.com/science/article/pii/S221313371...
I’ve used HDF5 before but not ASDF so I can’t fully evaluate their points one way or the other.
Adaptable Seismic Data Format: https://asdf-definition.readthedocs.io/en/latest/
Not sure how widespread YAML 1.2 adoption is though.
I agree YAML is bad and would say JSON is the best we have.
What are you suggesting? I think you have to make call if you say something is terrible or a bad choice.
(it's quite close to json, but supports comments and also integer sizes and such)
JSON is a great text-based exchange format and protocol medium, but extremely poor configuration language (no comments, no references, very restricted syntax for human editing).
Every tool has it uses (for configuration purposes I suggest something like HOCON, or even some superset of INI, like Python’s configparser).
YAML is bad at everything.
The Docker Compose and OpenAPI formats, are good uses of YAML that would be cumbersome in any other format.
Is HOCON better? Maybe. As far as I recall it doesn't have any affordance for multi-line strings, which I see as a valuable YAML feature. It does have its own merits, though, and is probability a better default than YAML in a lot of projects.
But at least if you want an INI-like format, use TOML instead of an ad-hoc underspecified alternative.
It isn't, really. A conforming YAML parser will not treat JSON data the same way that a JSON parser will in all cases. [0,1] The only correct way to deal with JSON is using a JSON parser.
But I'd love a shorter explanation!
"
The following lists our key requirements for a useful format:
• Human readable metadata; format suitable for archives.
• Efficient support for binary data.
• Implicit grouping and organization of metadata and data items.
• Use of a standard format for metadata to leverage existing tools and community.
• Easily extensible features, both for the general standard, and narrower needs.
• Strong validation tools.
• Ability to provide flexible WCS (world coordinate system) models.
• References to common data or metadata without requiring copying.
• Open Source, community controlled.
" Source : https://aspbooks.org/a/volumes/article_details/?paper_id=405...
These days, your storage-adjacent nodes can burn compute on the fly, you have to try hard to pay for rack-adjacent 10 gig, Snappy and zstd are so fast that SREs at Google and FB turn them on or off under service owners to shave the margins.
You can afford getting out of and back into a custom format for your domain, multiple times.
It’s going to be mmap’d ndarrays before it hits the compute stuff, and maybe libraries could be stronger there, but you can always throw away column names and putting them back is hard.
Now if I could stop twitching from the five years of astronomical ontology analysis, I would feel the world has truly moved on.
The timing of previous efforts was swamped by XML. Trying to maintain a clear distinction between "how" (the file format) and "what" (the content of a file in that format) can get lost in the quest to hang on to the "why" (the intent of the formatted data).
I'm an optimist. There's way more experience now. Consensus around Python has helped immensely in keeping this grounded with the scientists. I have hope.
squares: !core/ndarray-1.0.0
This is supposed to be a format, but a numpy array is a python concept. So that seems like a weird mismatch. And it is just a block of numbers really, so it seems odd to make it a special python thing.I also wasn't able to see how to store more complicated data, like a symmetric sparse matrix.
Are the other solutions not fast enough? If so what is the performance delta using ASDF? Are the other solutions not scalable enough? Likely not, given that ASDF mentions certain scale limitations-- but maybe those limitations are still less limiting than limitations of other solutions?
Whatever the reasons for making ADFS, it would be helpful to describe in the intro "here is a problem that ADFS solves" and "here is a graph (or similar) showing how it solves it better than alternatives".
Without that context, the description of Features, to my naive outsider's view, is not very valuable and I'm left still asking, "why use this instead of other solutions that many engineers already are familiar with?"
A directory on a file system happens to have these capability already. If you absolutely need everything in a single enormous file perhaps consider using Sqlite.
Python already has great support for these formats with Dask and Xarray. (Think multidimensional Pandas.)
Of course solving the general problem of all ROOT I/O (which can involve custom serialization or schema migration code in C++) won't (and doesn't need to, for most people) work.
And that is the exact usecase here.
The closest discussion on the format's raison d'être is here -- "Introduction: why another data format?" - https://www.sciencedirect.com/science/article/pii/S221313371... ... which in turn points people to other papers on this, such as https://www.sciencedirect.com/science/article/pii/S221313371... .
The authors say "There are a few worth considering. We will briefly review the landscape and comment on them. In short, we find significant problems with all of them. If one were to choose the best of them, it would likely be HDF5." .. and go on to list these "significant" problems as --
1. It is an entirely binary format.
2. It is not "self documenting".
3. There is only one implementation due to the complexity.
4. HDF5 does not lend itself to supporting simpler, smaller text-based data files.
5. The HDF5 Abstract Data Model is not flexible enough to represent the structures we need to represent, notably for generalized WCS.
The last seems the most significant to me. Let's see what WCS is about -- https://www.sciencedirect.com/science/article/pii/S221313371...
"WCS objects consist of a sequence of coordinate frames, with a transform definition from one to the next."
.. and they go on to give a "complex" example - https://www.sciencedirect.com/science/article/pii/S221313371...
I'm tempted to reference https://xkcd.com/927/ , but there is nothing wrong about having a specific rich file format for astronomy data. It is just that I think folks coming to the format ought to be told that up front.
No?