Feather: A Fast On-Disk Format for Data Frames for R and Python
blog.rstudio.org
blog.rstudio.org
What's the advantage of Feather over HDF5? Couldn't the Feather libraries be written with the same API but HDF5 as the storage format, if the Feather API is preferable?
So, python packages like h5py do not even try to release the GIL.
This makes working with HDF5 very annoying in python (when using multiple threads).
HDF5 is a really great piece of software -- I wrote the first implementation of pandas's HDF5 integration (pandas.HDFStore) and Jeff Reback really went to town building out functionality and optimizing it for many different use cases.
But the HDF5 C libraries are very heavy dependency. Feather by comparison is an extremely small amount of code (< 2KLOC in the core library) and a correspondingly minimal API. It's a simple file format with excellent performance, and we wanted to make it as easy as possible for people to use Feather.
There is also the Apache Arrow factor -- integration between the Arrow memory representation and R and Python tools will have a lot of ecosystem benefits, so one of the goals of Feather is to reconcile Python's and R's metadata requirements with the "official" Arrow metadata so that we can move around data frames with very low overhead.
I’d love to see a nice format of this type that can easily be written/read from Javascript in a browser [e.g. to get the data into a D3 visualization] and from Matlab, in addition to Python.
I looked into trying to implement an HDF5 codec in Javascript, but that looked like a large task for one person unfamiliar with the format.
Edit: https://github.com/wesm/feather/blob/master/doc/FORMAT.md
Seems a bit sparse/incomplete still (as would be expected for a brand new project).
My biggest personal annoyance is that HDF5 isn't thread safe^, so it only supports parallel reading and writing via multiple processes. This makes parallel computing a pain.
This is especially annoying when using HDF5's built-in compression, which hogs a lot of CPU. Inter-process communication is slower than reading from SSDs, so that isn't a great alternative: http://matthewrocklin.com/blog/work/2015/12/29/data-bandwidt...
There's a lot to be said for file formats that you can simply memory map, and that's exactly what Feather/Arrow are. Building out-of-core workflows on top of should be a joy.
Wes -- does the Python library for Feather already release the GIL?
^ you can use and/or compile HDF5 with a global lock, but the underlying library still isn't thread safe.
Checking out the code: https://github.com/wesm/feather/blob/master/python/feather/l...
So the actual reading from and writing to disk should be with a released GIL if I'm reading this correctly. The conversion to and from arrays or dataframes holds the GIL.
EDIT: What I DON'T know is how much libdataframe would look than libsqlite.
Is the feather in insights from feather the right word? It reads awkwardly to me, which could just be me lacking context.
- Both R and Python support strings, factors, and complex objects in a dataframe. What is NOT supported by feather?
- Feather is "not for long term data storage". Will it be standardize in a distant future?
- Do you plan to integrate it into Pandas?
I have no plans to integrate it with pandas, but I'm sure Wes does ;)
This appears to be "only" a serialization format ("oh, my unicorn only lays golden eggs"). I really hope this is the start of some common library infrastructure that can be used for all aspects of in- and out-of-memory data frames.
Great work, and I hope it is a harbinger of good things to come. Also, I'll treat this as tangible evidence that the "language wars" are stupid.
elif [[ "$OSTYPE" == "cygwin"* ]]; then
PARALLEL=$NUMBER_OF_PROCESSORS
And you'll need to use msbuild rather than make, e.g.: if [[ "$OSTYPE" == "cygwin"* ]]; then
msbuild gtest.sln /p:configuration=release
That got the 3rd party stuff working. But then I hit a snag, because building python 2.7 modules on Windows requires an old MSVC version that doesn't support stdint.h, which is used by feather in ext.cpp . Maybe a simple conditional compilation for the appropriate header will be enough to fix that, but I haven't got time to check today. So hopefully someone else can fix that...Is this like BigMemory but for data frames?
https://cran.r-project.org/web/packages/bigmemory/index.html
Thanks.
Either way, the idea of mixed Python/R pipelines with feather file intermediates input/outputs is pretty sweet. Learn in scikit, save to feather, plot in ggplot2... using Make to tie the pieces together?
(I'll probably take a stab at writing one this weekend, though)
Hadley, Wes, what are your thoughts on how to implement compression? I recall some open source columnar datastores (e.g. infobright) that achieved very VERY fast compression rates with just a few tricks: https://news.ycombinator.com/item?id=8354416
In particular, compression is extremely fast for columnar datastores (its the same type one after the other). Since a lot of times the data is sorted by some ID (date, individual, etc.), you should see large improvements in both speed and disk space.
Incidentally I had filed a bug request for a functionality to save the entire workspace in Pandas...but was rejected as being unpythonic. Oh and the devs claimed Apache Arrow was vaporware !
Is it stemming from a fundamental aspect of the data format - for example can you save two data frames to the same file?
Because if you can save two - why not save two hundred.
If not this, then I pray for Feather to be able to save multiple data frames innone file.
If you're looking for a container to store lots of tabular data in one file, I'd suggest SQLite. Using dplyr, you can save those dataframes very easily. Plus, you can join tables and perform efficient aggregations on datasets too large to keep in memory.
In a lot of ways, I don't understand what limitations prevent SQLite from becoming the defacto common data.frame format. There probably are some, I just don't understand the tradeoffs (especially given how much SQLite gives you for free)!
But coming back to the anti-pattern : well, obviously the authors have the power to not spend time on something. But I'm trying to figure out why it's an anti-pattern in general. Snapshotting execution state is probably the ideal goal, but saving intermediate data structures is a decent convenience feature.
Now if that's restricted by the limitations of the format itself (no multiple frames in a single file), then we are back to thinking that HDF5/sqlite may indeed be the better format.
It is convenient to save your complete workspace but I've seen too many cases where it's contributed to lack of reproducibility to spend my time working on it.
So I snapshot the workspace after I do a run and do some experiments. Now - for me, saving the workspace is a convenience feature, NOT a programming feature.
This is what I mean by thou-shall-not. My use case is very well defined and I'm not stupid. And I completely knows the pitfalls of what you talk about - but a philosophical opposition is what hurts me (and lots of devs like me)
(And even for your use case I would think you'd be better off keeping the models in a list and saving that. Then other random stuff in your evn won't get carried along for the ride)
Having said that I'm pretty excited about feather and where development will lead. I use RDS files quite heavily, mostly because of the compression which allows much smaller file sizes for distribution.[2] However, there's a trade off with parsing spread and also interoperability between languages. Looks like feather already has the interoperability part sorted, just waiting for compression now. Reading slices direct from disk is pretty exciting too.
[1] http://www.starkingdom.co.uk/faster-csv-import-with-r/
[2] http://www.starkingdom.co.uk/faster-import-with-r-redux/
Of course it all depends on what you call 'slow'. Reading a few hundreds of megabytes of megabytes of csv's isn't going to be 'really' slow on modern hardware even if it was fgetc'd character by character.
Either way: anything that represents data as text will become a bottleneck when the size of the dataset grows.
Relative to what?
The links I provided show that, using fread from data.table in R, parsing CSV data can be quicker than returning the same data from, for example, an SQLite database.
I believe the links do show CSV parsing relative to a couple of other common data parsing methods in R. Whether they're all "slow methods" is an entirely different question.
If you'd like to compare parsing CSV data relative to a "fast method", I'd very much like to read the analysis.
My point was, all conversions from text into numbers (I'm assuming here that the target use case is reading large amounts of numeric data) are slow and the slow part isn't the IO, but the conversion from text to numbers. In that light, it isn't a surprise that sqlite isn't much faster than csv, because sqlite doesn't have 'strong typing' itself. Any storage format that is concerned with speed will store data in binary format. But of course text formats are a lot easier to work with, and to interface between programs. I sometimes (when I have a lot of data that I know need multiple passes of reading) build quick and dirty 'caches' where I read file.csv and do what is basically a memory dump of the parsed data into file.csv.bin. My read functions can then check if that file exists and skip the parsing step. In my experience, this can be easily an order of magnitude (10x) faster. It's not portable or even elegant, of course.
Apart from that - when speed is of importance, one wouldn't use R in the first place, of course (I say that as someone who likes R for what it is and for the things it does well).
I don't do write ups of speed analyses, nor do I know of any, so I have to cop out on that one.
I don't know about R, but in Python, operations on objects such as NumPy arrays and Pandas DataFrames are all implemented using fast C code, and so is Feather. You can be concerned about speed.
Of course in-memory operations can be implemented efficiently, R does that too. But Python needs to parse CSV into numbers just like everybody else, and even if it's done in C underneath, it'll still be 'slow' (for some values of that word).
Relative to HDF5, and relative to Feather if it's doing its job right.
By comparison, Feather performs very close to disk performance. So speeds exceeding 500 MB/s (versus < 100MB/s for CSV) are common.