What's the advantage of Feather over HDF5? Couldn't the Feather libraries be written with the same API but HDF5 as the storage format, if the Feather API is preferable?
What's the advantage of Feather over HDF5? Couldn't the Feather libraries be written with the same API but HDF5 as the storage format, if the Feather API is preferable?
HDF5 is a really great piece of software -- I wrote the first implementation of pandas's HDF5 integration (pandas.HDFStore) and Jeff Reback really went to town building out functionality and optimizing it for many different use cases.
But the HDF5 C libraries are very heavy dependency. Feather by comparison is an extremely small amount of code (< 2KLOC in the core library) and a correspondingly minimal API. It's a simple file format with excellent performance, and we wanted to make it as easy as possible for people to use Feather.
There is also the Apache Arrow factor -- integration between the Arrow memory representation and R and Python tools will have a lot of ecosystem benefits, so one of the goals of Feather is to reconcile Python's and R's metadata requirements with the "official" Arrow metadata so that we can move around data frames with very low overhead.
I’d love to see a nice format of this type that can easily be written/read from Javascript in a browser [e.g. to get the data into a D3 visualization] and from Matlab, in addition to Python.
I looked into trying to implement an HDF5 codec in Javascript, but that looked like a large task for one person unfamiliar with the format.
Edit: https://github.com/wesm/feather/blob/master/doc/FORMAT.md
Seems a bit sparse/incomplete still (as would be expected for a brand new project).
My biggest personal annoyance is that HDF5 isn't thread safe^, so it only supports parallel reading and writing via multiple processes. This makes parallel computing a pain.
This is especially annoying when using HDF5's built-in compression, which hogs a lot of CPU. Inter-process communication is slower than reading from SSDs, so that isn't a great alternative: http://matthewrocklin.com/blog/work/2015/12/29/data-bandwidt...
There's a lot to be said for file formats that you can simply memory map, and that's exactly what Feather/Arrow are. Building out-of-core workflows on top of should be a joy.
Wes -- does the Python library for Feather already release the GIL?
^ you can use and/or compile HDF5 with a global lock, but the underlying library still isn't thread safe.
Checking out the code: https://github.com/wesm/feather/blob/master/python/feather/l...
So the actual reading from and writing to disk should be with a released GIL if I'm reading this correctly. The conversion to and from arrays or dataframes holds the GIL.
So, python packages like h5py do not even try to release the GIL.
This makes working with HDF5 very annoying in python (when using multiple threads).