Pandas 0.15 has been released
pandas.pydata.org
pandas.pydata.org
I used to program solely in R, but after discovering pandas I really have no need to go back to R. My project workflow consists of several IPython notebooks+pandas+sklearn.
Works extremely well in production, as in, on a flask web server, as well.
For other applications (especially charting and data manipulation with ggplot2 and dplyr respectively), R has an edge.
Seaborn - http://web.stanford.edu/~mwaskom/software/seaborn/
Blaze - http://blaze.pydata.org/docs/v_0_6_5/index.html
Bokeh - http://bokeh.pydata.org/
If I had stronger python-fu I would love to build "GGPy".
www.plot.ly
Categorical is awesome, and equivalent to R's c() iirc. This should make plots easier in terms of automatically deciding whether to facet something, or showing legends nicely etc.
The memory usage feature is super neat.
Also for those of us stuck with STATA, the to_stata() and read_stata() just got much better.
I'm eagerly awaiting a numpy native NA value instead of np.NaN.
There are no links in docs to the types that are being referenced, some types do not have documentation, some documentation is the function header with no other info ie no documentation, functions that take string formatting info eg '5min' do not have their argument possibilities documented anywhere I can find.
The argument possibilities has always been an issue for me. In general, I have found, if you have non-homogenous data, Pandas is your best bet due to how general it is, even if it is sometimes frustrating when it forces a generalization on your data (e.g., try doing type conversions on numpy.datetime64, it lacks any sort of intuition).
The distinction between a Series and a dataframe is something I always found pretty silly/frustrating and I wonder if it was the result of an early implementation issue rather than a logical simplification.
maybe I'm particularly ignorant of some issues since my use case is perhaps more straightforward than some others but I've built an entire labelled data library for my team and it is easier for us to operate under similar primitive beliefs to the numpy ndarray, i.e., it's always an ndarray, adding a column (or dimension) does not change the type of the object and the associated methods and indexing in one particular way vs another (df['a'] vs df[['a']]) does not change the type of your object.
If I'm missing the point of Series I would love to see them justified or a use case referenced.
FWIW, there are over 1500 pages in the docs including a short-over view, tutorials, extensive feature coverage, and interaction with other tools: http://pandas.pydata.org/pandas-docs/version/0.15.0/pandas.p...
The docs may have some issues, but they certainly can't be characterized as lacking.
Don't get me wrong, I'm grateful for all the work and I know I haven't contributed much but I think the online could be improved with more examples and recipes.
Edit: There really is no excuse, getting started is easy[2].
The issue I had was not the documentation but the language of pandas mirrors the language used in R (I think this is something Wes McKinney intentional did) and it's the burden of all that new verbage that makes the documentation harder to sift through. Some choice exampels; "melt", "stack/unstack" and "reindex" — necessary, I grant you, so that functions can be aptly named and in turn encapsulate vectorised procedures that are composable.
I found that the documentation was harder to search because I lacked the domain language and the documentation, for better for worse, doesn't dawdle with educating the reader about the verbage — worked examples often provide a easier route. It reads like a mathematical proof rather than prose and I used to think that the documentation was too terse but now I appreciate that probably just succinct.
https://github.com/pydata/pandas/issues/7517
As @mynegation notes, you can use Anaconda (or virtualenv).
This is the real killer. Any leads on making it less closed?
Also, I'm unfamiliar with how pip works, but you can't even install into your local user profile (i.e. w/o root)?
In particular, note that NumPy 1.7.0 or newer is required.