The structure of a HDF5 file can be manipulated to put data where you want it in a file. For many files of the same type, you can set the metablock size so that your datasets are always at the same offsets.
You can use tools like h5ls or h5dump to describe the structure of a HDF5 file. It turns out that if you do this, you can make HDF5 cloud friendly.
https://medium.com/pangeo/cloud-performant-netcdf4-hdf5-with...
String data and HDF5 = No.
It does have an index tho compared to Parquet, which is handy for ML type stuff.
It's a quantum file format. It works so long as you aren't looking at it too much.
[1] https://numpy.org/devdocs/reference/generated/numpy.lib.form...
Nympy has a reasonably sensible format, but I haven’t tried to do anything tricky with it.
An online, database-like product is most important if you need to coordinate concurrent activity by a number of data producers and consumers while maintaining a coherent view of the growing data. If you can break it into distinct phases with individual actors, passive serialization formats can make more sense. Adoption of bject-storage semantics would help eliminate some of the corruption/concurrency hazards mentioned in Cyrille's post. You write entire files and expose them to read-only consumers once they are valid and complete, side-stepping concurrent writer/reader scenarios.
However, object-storage still has coherence problems. If you expect metadata to need rounds of editing or curation, you don't want it embedded in bulk objects, where mutation is expensive, i.e. rewriting an entirely new version of the object. It is easier with a companion file strategy, where you can regenerate smaller metadata files alongside immutable bulk data files. The object store can ensure that concurrent users are only encountering a coherent snapshot of each metadata file, i.e. before or after it was replaced. But it does not provide coherent views of multi-file changes. You avoid low-level codec failures from concurrent writes to shared files, but you may still have semantic inconsistencies when interpreting the evolving multi-object dataset.
A hybrid approach would be to have some other data-management system with database-like properties to track the objects. The database stores a coherent catalog of which data objects are available. New objects are added and registered in the catalog for others to read, but you don't mutate them afterwards. The catalog should contain metadata needed to coordinate the workflows of production and consumption of objects. You could benefit from using the catalog as an authoritative store of metadata that is still being curated/refined. But you can also export snapshots of metadata into companion objects at appropriate milestones in the processing workflow. These could even be versioned to record well-defined "data releases" or provenance info related to individual processing tasks.
Of course, this still leaves the question of what file format(s) to use for the individual objects in the system... one could do all of the above and still choose to use HDF5! But these individual files would be simpler, i.e. to store a single N-dimensional numerical array in one file. Microscopy projects might choose OME-TIFF for individual image stacks, or something even more mundane like a directory of individual 2D TIFF or PNG files, if that maps well to how their acquisition or processing pipeline works. Sometimes, the right file layout is just as important as choosing chunking parameters within a format like HDF5. Naive pipelines often process whole files in RAM, so you end up choosing appropriate file layouts to align to the kind of sparse or sequential access needed by data producer and consumer programs.
A file is just an abstraction over a block device. HDF5 is a meta abstraction, storing multiple file abstractions within an existing container file abstraction. There is nothing conceptually wrong about this but incredible to me that anyone would have invested the time into creating an abstraction that looks exactly like its container just to avoid having to tar it.
There are indeed a lot of reasons not to use HDF5: if you have a lot of small records, are primarily storing strings, or you don't share your data with many colleagues, there are a lot of better alternatives. But if you just want to dump a few hundred GB into one massive file and have decent IO in a format people can figure out, it works pretty good.
If you target Linux, those work out of the box due to the file system doing the lifting.
With windows we had to use the short file name (e.g. FILEN~1.EXT) as a workaround.
Also, you have to watch out for what the library does when it writes to (Edit: Windows) network shares - we've seen it write NaNs where there were values in the data it should save, maybe a latency issue, maybe a configuration issue - but not something you get told about.
I would like to move away from it.
This is definitely not true, at least for C++, but I agree that it can be hassle.