Pandas 0.7.0 released: Python data analysis library
pandas.pydata.org
pandas.pydata.org
http://pandas.pydata.org/pandas-docs/dev/io.html#hdf5-pytabl...
Together with scikits.learn, this could prove really useful in machine learning and data analysis projects.
That's just wrong.
Conflating the two concepts means you can't tell the difference given the result set. It's just a happy accident that "unknown" and NaN have identical propagation rules, but that doesn't mean that it's safe to use one in place of the other. Reading up on it, it looks like Octave and Matlab can treat NaN as "missing data", though, so I guess there's a certain "industry standard behaviour" to follow so as not to surprise users, but it's still less than ideal.
In an ideal world, we could define an explicit "missing data" quiet NaN which would have a distinct visual representation - I suspect this is doable with access to the float exponent bits, but I don't know how Python could take advantage of it.
I do agree that NaN's are a better choice for truly missing data, but I'm biased just because they use less memory. They're not a solution for non-floating point data, though.
Great job on Pandas, by the way!
FYI it's fun to hear an academic ragging on "unmaintainable code".
One of the strengths of Python is that you can use it to build critical production systems (which I've done for many years in the financial industry). You come up against a lot of people who think "Java/C++/C# are the only suitable systems languages".
I'm merely pointing out that the language bashing is not productive. The writeup should point out the positives and stop trying to turn the differences between languages into a parallel of state of American political discourse.