Pandas 0.4 (Python data analysis library) released
pandas.sourceforge.net
pandas.sourceforge.net
I wonder what the motivation is to do this when R is so mature (especially in the availability of specialized packages), and available through RPy.
The bigger picture reason "why not R" is that R is not very suitable for building production systems. I started building this library while working for AQR, a quant hedge fund, and needed to have statistical computing building blocks integrated with a much larger system. R is a mediocre programming language and has very weak general purpose libraries. But amazingly good data visualization and mature statistics libraries indeed. Using R as a black box (e.g. via RPy, Rcpp, or RJava) is a good idea in theory, but recovering from and dealing with errors/exceptions with real world data is a very thorny problem. Plus maintaining a big pile of R code is kind of a nightmare (believe me, been there, done that!).
Looking forward to checking this out.
It would be an interesting avenue to pursue building a "big data" on-disk OLAP engine with pandas-like semantics (e.g. expressing groupby operations with the same syntax but operating on big data on disk or across a cluster of computation nodes).
x y z value 0 0 0 32 0 0 1 64 0 1 0 23 0 1 1 3.14 1 0 0 4.3 etc.
This is essentially a 3D cube of data, no hierarchical indexing involved. The benefit of hierarchical indexing is that you can wrap your spatial dimensions into a single real dimension for e.g. code abstraction.
I have actually been developing a similar library for OCaml (even with hierarchical indexing). It is good to see our libraries share many of the same ideas! I wonder though, have you considered GPU acceleration? AFAIK neither Matlab nor R do this natively yet.
http://www.meetup.com/nyhackr/events/28880161/?hidePromoBar=...