Moving away from HDF5
cyrille.rossant.net
cyrille.rossant.net
Luckily the authors of Alembic were smart and in their initial design abstracted out the HDF5 interface and were able to provide an alternative IO layer based on C++ STL streams. The C++ STL streams-based interface greatly outperformed the HDF5 layer.
Details on that transition here:
https://groups.google.com/forum/#!msg/alembic-discussion/FTG...
Although the spec is huge, there is plenty of sample code online to get it working. You do actually have to read it though to understand slabs, hyperslabs, strides etc. Once you get though, its really versitle.
As far as speed, we used it to replace our propriatary data format. We would have to provide readers to all the scientists that use our data. It was nightmare. Some people want stuff in R, some in Python 2.7, some in Python 3.4, some in Matlab, and the list goes on. HDF5 gets rid of all this.
When in the field and the system shits the bed, its really easy to open an HDF5 file in HDFview and inspect the file contents. I dont always have matlab available when im in the field, same with python. Sometimes I just need to look at a time series and I can diagnose the problems with the system.
For me, its silly in 2016 to have any kind of proprietary binary format when something HDF5 exists.
Many of the complaints the author had makes me think he the stereotypical scientist of really smart in one area but cant program worth the beans. I dont think that HDF5's fault.
"When in the field and the system shits the bed", to quote you... you can just open the dataset in Windows Explorer. Or Mac OS Finder. Or Nautilus. Or using 'cd' and 'ls' in Linux. Want to look at an array? Sure, open it in a hex editor or Python or, heck, FORTRAN88. You can .tar up your folder (or subfolder, or any arbitrary subset of the dataset) and send it to someone with no knowledge of the format whatsoever, and they'll be able to make sense of it in minutes. This isn't anything remotely complex - it's just using the filesystem rather than creating a filesystem-within-a-filesystem. Want to keep track of changes in your massive dataset? Sure, just back it up on Mac OS X Time Machine, or a simple rsync script, or even Git LFS.
Most researchers aren't computer scientists; they know how to use Dropbox and Notepad++ and open text files, and they don't want to have to install a Java-based HDF5View when they could just use Windows Explorer.
It's not even that HDF5 is that bad, it's just that filesystems are, in many respects, so much better.
(If it wasn't clear from the article, we're not just misinformed - we're making this call after having developed an entire software suite around HDF5, spent about two years of firefighting HDF5 issues and wasted days of development time (so many horror stories) - this is actual feedback from several dozen users, thousands of datasets, and petabytes of data.)
HDFfView is not the only viewer in town. There are several viewers.
I've used HDF5 weekly for the last 5 years and so have my associates and its been wonderful. MATLAB even uses it to store its .mat files these days.
I think you are misinformed and Im sticking to my story !
It's rare we ever need the whole dataset - in fact, it's really great to be able to say to the user "don't send us your 100GB dataset: just go to the "acquisition" subfolder and send me the 10MB file called "oscilloscope.dat"". With HDF5 this is difficult enough that it's almost always easier to send 99.9% of useless data (i.e. the whole file) when all you want is a single array within it.
If your users will rarely need to do this, you could just store the entire folder hierarchy in a .zip and access it using standard tools that most programming languages have. It's worth noting that the new Microsoft Office formats do exactly this - in their case, a bunch of XML files inside a .ZIP. (Rename a .docx to .zip and you'll see!).
MATLAB has moved from their own custom binary format to HDF5, which is the lesser of two evils.
Yes, I do know about docx and zip.
I'm happy with what matlab has done. People send me .mat files and I happily process them in python. And my plots usually look much nicer also. :)
Are you not concerned about i-node consumption with the file-system based approach?
I kind of the like the approach suggested above - if you have many small files store them in a zip archive and use some library to access the data directly.
It's easy to get a file format working on one system. It's Herculean getting it working on all systems bug free. It's nearly impossible to get something to work portably and performant across many systems.
As for simplicity, people start wanting metadata and headers and this and that, and before you know it you need HDF5 or ROOT again and it's no longer simple. Maybe if you're lucky you can stick with something that looks like FITS. If it's tabular, SQLite still can't be beat. Maybe Parquet would work fine too.
I'd vehemently oppose anyone in the projects I work on from trying to standardize on a new in-house format. I'd maybe be okay if they were just building on top of MessagePack or Cap'n Proto/thrift etc... but nearly every disadvantage the OP references about HDF5 will undoubtedly be in anything they cook up themselves. For example, a "simpler format" that works well on distributed architectures, well... now you're going to go back to the single implementation problem.
We can get by just fine with: - N-dimensional arrays stored on disk - Key-value metadata associated with those arrays - A hierarchical data structure.
We've been very happy so far replacing HDF5 groups with folders (on the filesystem), HDF5 datasets with flat binary files stored on disk (just as HDF5/pretty much any other format stores them - each value takes up 1 or 2 or 4 bytes, and your filesize is just n_bytes_per_value * n_values), and attributes by JSON/XML/INI files. If I sent you one of our datasets, zipped up, you'd be able to make sense of it in a matter of minutes, even with no prior knowledge of how it was organised.
It is very tricky to build something that works reliably across all systems, but, thankfully, filesystem designers have done that job for us. And filesystems are now at a point where they're very good at storing arbitrary blobs of data (which wasn't the case when HDF was founded). Filesystem manipulation tools (Windows Explorer / Finder / cd/cat/ls/mkdir/mv/cp/[...]) are also very good and user-friendly.
There isn't really anything we miss about HDF5 at all. Perhaps if your project has spectacularly complex data storage requirements (as to your examples: metadata/headers are easily stored in JSON), but there's no other project I know of that actually relies on an HDF5-only feature and couldn't trivially use the filesystem instead.
My main qualm with it was the claim about 100x worse performance than just using numpy.memmap(). To the author's credit, he posted his benchmarking code so we could try it ourselves. (Much appreciated.) But as it turned out, there were problems with his benchmark. A fair comparison shows a mixed picture -- hdf5 is faster in some cases, and numpy.memmap is faster in other cases. (You can read my back-and-forth about the benchmarking code in the blog's comments.)
One minor complaint about presentation: Once the benchmarking claims were shown to be bogus, the author should have removed that section from the post, or added an inline "EDIT:" comment. Instead, he merely revised the text to remove any specific numbers, and he didn't add any inline text indicating that the post had been edited.
I think the rest of the post (without performance complaints) is strong enough to stand on its own. After all, performance isn't everything. In fact, I'd say it's a minor consideration compared to the other points.
When it comes to performance, I think the main issue is this: When you have to "roll your own" solution, you become intimately aware of the performance trade-offs you're making. HDF5 is so configurable and yet so opaque that it's tough to understand why you're not seeing the performance you expect.
...it's worth noting that many of the tricky things about tuning hdf5 performance are not unique to HDF5. For storing ND data, there will always be decisions to make about when to load data into RAM vs. accessing it on demand, whether or not to store the data in "chunks", what the size of those chunks should be (based on your anticipated access patterns), whether/how to compress the data, etc. These are generally hard problems; we can't blame HDF5 for all of them.
For simple uses I now use '%a' printf specifier. It is specifically designed to avoid losing a single bit of informaiton. And you can easily read floats stored this way in numpy by using genfromtxt with optional converters= argument and builtin float.fromhex function.
It looks like you can just use these functions to convert between endianness conditionally based on the platform if you are in C/C++:
http://stackoverflow.com/a/8671129
And boost has one here: http://www.boost.org/doc/libs/1_58_0/libs/endian/doc/index.h...
Your comment is true for your use case, but it's not really responsive to the technical issue here.
for (i=start; i<end; i++) arr[i] = ntohl(arr[i]);
That's not more code than calling printf/scanf.The standard does not specify endianness [1]. From the standard point of view float is a sequence of bits. But if you memory-map the file and store float on little-endian machine, you get bytes reversed compared to how it is written in the standard. On big-endian machine sign is stored in the first byte, and on little-endian machine sign is stored in the last byte.
edit: oh, the Java tooling was already mentioned.
[1] http://static.googleusercontent.com/media/research.google.co... [2] https://github.com/h5py/h5py/blob/master/h5py/_locks.pxi
The schema-based serialization libraries (Thrift, Protobufs) or MsgPack are a good way to avoid that too. They come with a lot less baggage than say HDF5. Also, if efficiency isn't paramount -- just use SQLite! Amazing tool when it's in its sweetspot.
Lots of tradeoffs when dealing w/ serialization and file formats, no easy answers.
its also really easy to write a viewer in python using matplotlib.
If you want headers, you can define those as text and just use extents to embed the binary data. Tar is another alternative, pretty easy to implement.
Research data are almost always specific. The idea of a general format serving all research just sounds silly to me.
HDF provides command-line tools like h5dump and h5diff, so you can dump HDF5 file to text and pipe it into standard and non-standard unix tools [1].
[1] https://www.hdfgroup.org/products/hdf5_tools/index.html#h5di...
http://aadrake.com/command-line-tools-can-be-235x-faster-tha...
Another thing that will happen is this: if you just use a directory full of files as a single logical file, you will end up writing code that does the equivalent of 'rm -rf ${somedir}' because when users choose to overwrite "files" (really, a directory), you need to clear out any previous data so experiment runs don't get mixed. You can easily see where this can go bad; you will have to take extraordinary care.
Even better directly use BTRFS (or ZFS) send / recv to send to move to another system with that file system.
Can also pipe tar directly to something like nc. Remember benchmarked moving a large number of files on a LAN and tar + nc came out as the winner over say rsync or scp. It saturated 1Gbps connection pretty close to its maximum expected capacity.
You'd obviously not expect user to do that by hand and would write tools to do it.
While it's true that you lose atomicity with this, it can both burn you and help you: you can track specific parts of your dataset in revision control, email subsets of it back and forth, combine datasets easily, or even store it across several servers and manage it with symlinks, for example.
Our users are aware of this and it isn't really a problem for our use-case. But if you're worried, there's always the option of having your whole dataset as a ZIP/TAR file (like all of Microsoft Office file formats are - XML files within a .ZIP); tools for modifying folders within ZIP are much more well-established than HDFView, and most modern programming languages provide libraries to read and modify files within archives without unzipping them; you could make your program agnostic to the files being within an archive (high portability, lower performance) or directly within the FS (loss of atomicity, easier/faster to use).
I ended up designing a transactional file format (a kind of log-structured file system with data versioning and garbage collection, all stored in a single file) from scratch to match our requirements. Works flawlessly with terabytes of data.
GC was offline though; it was performed at the time the container was "upgraded" from RO to RW access. I don't think it'd be difficult to make it online, but there was no need for that.
This is where they lost me. This is NOT a lot of data. Should we be surprised that memory-mapping works well here?
Below about 100-200 GB you can do everything in memory. You simply don't need fancy file-systems. These systems are for actual big data sets where you have several terabytes to several petabytes.
Don't try to use a chainsaw to cut a piece of paper and then complain that scissors work better. Of course they do...
Also, as I said, the 100,000 x 1000 example is a quite optimistic one, we do have cases now with 100,000,000 x 10,000 arrays, and this is only going to increase in the months to come with the new generation of devices.
Bugs and crashes in the HDF5 library and in the wrappers - sure, every sw has bugs. But in over 15 years of using HDF, I have not seen a bug that stopped me from doing what I want. And the HDF team is very responsive in fixing / suggesting work arounds.
Poor performance in some situations - yes & no. A well built library with a well designed application should approach posix performance. But HDF is not a simple file format, so expect some overhead.
Limited support for parallel access - Parallel HDF is one of the most, if not the top most, popular library for parallel IO. Parallel HDF also uses MPI. If your app is not MPI, you cant use Parallel HDF. If the "parallel access" refers to threading, HDF has a thread safe feature that you need to enable when building the code. If "parallel access" refers to access from multiple processes, then HDF is not the right file format to use. you could do it for read-only purposes but not write. again, not the right motivation to pick HDF
Impossibility to explore datasets with standard Unix/Windows tools - again, HDF is not a flat file, so how can one expect standard tools to read it? its like saying I would like to use standard tools to explore a custom binary file format I came up with. wrong expectations.
Hard dependence on a single implementation of the library - afaik there is only one implementation of the spec. the author seems to know this before deciding on HDF. Why is this an issue if its already known?
High complexity of the specification and the implementation -
Opacity of the development and slow reactivity of the development team - slow reactivity to what? HDF source is available so one can go fix / modify whatever they want.
seems the author picked HDF with wrong assumptions.
HDF serves a huge community that has specific requirements, one of which is preserving precision, portability, parallel access, being able to read/write datasets, query the existing file for information of the data in the file, multi dimensional datasets, large amount of data to fit in a single file, etc.
I think it's reasonable to be very upset if you have a container file and adding new named chunks to the file has the possibility of causing the old data to become unreadable. It's fair in a crash before the file was saved that new chunks might be bad, but old chunks should be fine.
Our method is to make a copy of the file with a .tmp extension, make your mods, then rename the file and delete the old one.
HDF5 was never intended be be a container used to stream data into -- it was meant for sharing data.
http://cscads.rice.edu/workshops/summer-2012/slides/datavis/... Extreme IO scaling with HDF5
http://algoholic.eu/sec2j-journalling-for-hdf5/ HDF5 with a journal