HDF5eis: Storage IO solution for big multidimensional time series sensor data
pubs.geoscienceworld.org
pubs.geoscienceworld.org
I can't imagine there'll be a queue of people wanting to implement that for fun on their weekend
HDF5 is on the way out anyway.
On a more serious note, DuckDB is actually a pretty fantastic replacement for SQLite for some data analytics tasks. Hdf5 is still great for numerical data in this space though. I would welcome other options like the SQLite/DuckDB scenario for matrix data.
:D
The structure of a HDF5 file can be manipulated to put data where you want it in a file. For many files of the same type, you can set the metablock size so that your datasets are always at the same offsets.
You can use tools like h5ls or h5dump to describe the structure of a HDF5 file. It turns out that if you do this, you can make HDF5 cloud friendly.
https://medium.com/pangeo/cloud-performant-netcdf4-hdf5-with...
String data and HDF5 = No.
It does have an index tho compared to Parquet, which is handy for ML type stuff.
It's a quantum file format. It works so long as you aren't looking at it too much.
[1] https://numpy.org/devdocs/reference/generated/numpy.lib.form...
Nympy has a reasonably sensible format, but I haven’t tried to do anything tricky with it.
An online, database-like product is most important if you need to coordinate concurrent activity by a number of data producers and consumers while maintaining a coherent view of the growing data. If you can break it into distinct phases with individual actors, passive serialization formats can make more sense. Adoption of bject-storage semantics would help eliminate some of the corruption/concurrency hazards mentioned in Cyrille's post. You write entire files and expose them to read-only consumers once they are valid and complete, side-stepping concurrent writer/reader scenarios.
However, object-storage still has coherence problems. If you expect metadata to need rounds of editing or curation, you don't want it embedded in bulk objects, where mutation is expensive, i.e. rewriting an entirely new version of the object. It is easier with a companion file strategy, where you can regenerate smaller metadata files alongside immutable bulk data files. The object store can ensure that concurrent users are only encountering a coherent snapshot of each metadata file, i.e. before or after it was replaced. But it does not provide coherent views of multi-file changes. You avoid low-level codec failures from concurrent writes to shared files, but you may still have semantic inconsistencies when interpreting the evolving multi-object dataset.
A hybrid approach would be to have some other data-management system with database-like properties to track the objects. The database stores a coherent catalog of which data objects are available. New objects are added and registered in the catalog for others to read, but you don't mutate them afterwards. The catalog should contain metadata needed to coordinate the workflows of production and consumption of objects. You could benefit from using the catalog as an authoritative store of metadata that is still being curated/refined. But you can also export snapshots of metadata into companion objects at appropriate milestones in the processing workflow. These could even be versioned to record well-defined "data releases" or provenance info related to individual processing tasks.
Of course, this still leaves the question of what file format(s) to use for the individual objects in the system... one could do all of the above and still choose to use HDF5! But these individual files would be simpler, i.e. to store a single N-dimensional numerical array in one file. Microscopy projects might choose OME-TIFF for individual image stacks, or something even more mundane like a directory of individual 2D TIFF or PNG files, if that maps well to how their acquisition or processing pipeline works. Sometimes, the right file layout is just as important as choosing chunking parameters within a format like HDF5. Naive pipelines often process whole files in RAM, so you end up choosing appropriate file layouts to align to the kind of sparse or sequential access needed by data producer and consumer programs.
A file is just an abstraction over a block device. HDF5 is a meta abstraction, storing multiple file abstractions within an existing container file abstraction. There is nothing conceptually wrong about this but incredible to me that anyone would have invested the time into creating an abstraction that looks exactly like its container just to avoid having to tar it.
There are indeed a lot of reasons not to use HDF5: if you have a lot of small records, are primarily storing strings, or you don't share your data with many colleagues, there are a lot of better alternatives. But if you just want to dump a few hundred GB into one massive file and have decent IO in a format people can figure out, it works pretty good.
If you target Linux, those work out of the box due to the file system doing the lifting.
With windows we had to use the short file name (e.g. FILEN~1.EXT) as a workaround.
Also, you have to watch out for what the library does when it writes to (Edit: Windows) network shares - we've seen it write NaNs where there were values in the data it should save, maybe a latency issue, maybe a configuration issue - but not something you get told about.
I would like to move away from it.
This is definitely not true, at least for C++, but I agree that it can be hassle.
https://www.researchgate.net/publication/367373175_HDF5eis_A...
https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2021...
The fundamental problem with time series data is often that the insert pattern is the exact opposite of typical retrieval patterns. For example, you may insert 1000s of properties at once (most efficiently stored as interleaved data), yet typical access patterns involve obtaining a single property over many timestamps. Of course, it can work fine despite the inefficiency. If you have a few 100 sensors and are sampling them every minute, storing data for a month, it would likely be fine. If you plan on storing audio samples as records in a DB (i.e. one record per sample), it will fail.
As another commenter mentions, HDF5 doesn't really help much of this. It is an unnecessary rabbit hole to some extent. It mostly just a glitchy b-tree implementation with no WAL log + some compression algorithms.
ClickHouse is a relational DBMS, and it works for time series better than specialized systems (like TimescaleDB or InfluxDB). You can quickly get trillions of records for time series, but it is not a problem. For example, see https://www.youtube.com/watch?v=JlcI2Vfz_uk
If you try TimescaleDB in this scenario, it will barely be crawling.
PS. I am developer of ClickHouse, and I use it for these scenarios every time.
What does work for large regularly sampled datasets is array storage solutions like hdf, or image formats, etc etc.
It's the same reason and same approach for your vacation photos. You wouldn't put each pixel in the db, but you might very well have a db with a row for each photo.
Use a db for the non-regular portions (e.g. an index of datasets) and array storage for the actual data itself.
It's obviously a bit more work than relying on HDF5 to do it for you, but not that bad since you don't need to implement a full general solution.
A modern alternative would be Zarr[0] which was frustrated with the complexity of HDF5 and wanted something that would be more amenable to working with S3 storage. Never worked with it, but the ideas are laudable. Then again, I was never all that frustrated with HDF5. I wish the exploratory tooling was better, but that is about the end of my complaints.
There’s some cases where data really needs to be structured, especially closer to instruments producing data. The thing is - some or all of that data is going to end up in a database or a dataframe sooner or later
https://pypi.org/project/h5py/
If you're using another language there might be a more appropriate high level library.
Thanks
It also has a spec, which can be useful.