Modern Pandas (Part 2): Method Chaining
tomaugspurger.github.io
tomaugspurger.github.io
Sometime I wish designers of Numpy or scikit-learn should have developed Pandas.
As this is the top comment, can you (and others) at least post the problems so that we can have an intellectual discussion? Maybe the Pandas devs might take a point or two.
It's basically a DSL constructed out of the dismembered syntactic bones of Python, which breaks every piece of semantics in the host language that it possibly can. I'm sure this is convenient (and maybe even tractable to use) in an interactive notebook, where you can try out and verify behavior in real-time by looking at the output. But in any kind of non-immediate-feedback scenario where you're trying to engineer a production system, it's a hellscape of jumping back and forth to the documentation because even your most fundamental assumptions about the host language's semantics have been thrown out the window. On top of that (and largely because of it), it also resists static typing like crazy.
It also has a bunch of APIs where it's really hard to know what will and won't mutate, and others that are just made generally hard to wrap your brain around for the sake of saving a few characters.
Finally: some portion of the hate it gets is probably from engineers without a statistics background who aren't familiar with the huge dictionary of jargon and abbreviations its APIs use. This complaint is maybe less valid, since it is a data science library, but it certainly colors emotions and makes it even harder to deal with in a production context for lots of us. There were several times where, once I did some research and dug past the jargon and learned all the background for what it meant, the concept itself wasn't complicated. But Pandas did use the jargon, so I couldn't just understand, I had to go learn stats first and traverse all this needless indirection. It's an API that isn't designed for engineers but engineers often have to deal with it anyway.
Pandas' jargon isn't really even statistics jargon. For example data frames are tables, which are inherently rows and columns. Columns have names, rows can have names too (although that's not really needed). Pandas' documentation has very few mentions of rows and columns at all, instead they use "axes", "labels" and "index" and there's not even a proper explanation in documentation of what those mean. And that "index" is something that the user needs to manage, sometimes "drop", sometimes "reset" without really understanding why. It seems to me two things: (1) really bad choice of naming things and (2) exposing to users some performance-related details that should have been kept under the hood, if they are needed at all.
As for your other complaints, I would need to see examples because I'm not really sure what you mean. Are you referring to things like "mad" which stands for Mean Absolute Deviation? This I would put into my "grok documentation" bucket because just look up the mad method and it literally gives you the acronym. Would the method "mean_absolute_deviation" really help you? Furthermore, if they used long names like that, statisticians from R (which are the target audience) would have another nothing-burger bullet to use against Pandas.
Duckdb can run over Pandas dataframe and it is super fast. I haven't tried it yet. I am more familiar with Pandas than SQL so I would like to hear your opinion after you give it a shot.
I think it is a bit mean to say about a package as popular as this to have “no design philosophy”. You should read about their design philosophy before making that comment.
From my experience, I jumped all in when I discovered pandas and then I dialed it back. (It was partly because of my inexperience before.)
Pandas is more useful for exploratory data analysis. It is kind of in the philosophy of working in the terminal (with UNIX pipes, etc.) to explore things quickly. That’s why you’d see in the wild people chaining tons of methods together, sort of like people writing terminal one liner chaining a lot of pipes.
It is also useful as a dictionary containers, in fact you can treat a data frame as if it is a dictionary of dictionary of values in terms of API. Vice versa, if one has an internal structure that is dict of dict of values, you can convert that to a DataFrame as a drop in replacement (I’ve done that when working with a software that does not use pandas.)
For simple things that one has prototyped, it can be left as is for “production”.
But for more complicated things, one should “productionize” it using easier to understand and/or more performant logic.
Some of the mistakes of using pandas is to treat it as you “data container”, as if the table itself is self explanatory. From my experience I’ve been confused by the table I saved in the past. So now I write classes that has a to_frame method that my internal data structure can be converted to a dataframe for further exploration if needed.
The problem is nobody is using them.
From https://docs.python.org/3/howto/functional.html it's got map/filter/currying and plenty more, what's misting in your view?
nothing intrinsically 'functional' about this, though it's nice
> syntax for partial application
<https://docs.python.org/3/library/functools.html#functools.p...>
> and composition,
Hmm, to my surprise I can't find this but TBH it's trivial to write, here's an eg <https://stackoverflow.com/questions/16739290/composing-funct...>
> a typing system which can express structural types,
typing is orthogonal to FP (though very nice to have)
> a generalised list comprehension
IDK what this means, what's ungeneral about comprehensions currently?
Pythons failure to provide this, indeed, outright hostility to doing so is half the reason pandas is a mess of incomprehensible syntax
By prioritising assignment expressions and implementing pattern "matching" as a statement, they're clearly showing a lot of hostility to one of the major use cases of their language
(incidentally, recall that C# introduced LINQ in 2008, a generalised comprehension! That's C#.)
There isnt the syntax to reimplement to support natural phrasings of data transformation, in lieu of this, pandas exploits weird operators such as `.loc[a,b]`.
At some point the dam is going to have to break, and python is going to have to introduce something to resolve this mess. However, i'd bet it'll be a decade of tooth-and-nail fighting about it. It isnt a software engineering language for education any more, and i'm not seeing this reality being acknowledged
Commenters here are keen to say "but dont you know!" -- and, yes, I do.
Python does FP fine and FP via an ugly library or via elegant inbuilt syntax is still FP. I get you want nicer syntax but it wasn't clear.
I dunno, write a PEP and see if it gets support. I hope you succeed!
Available in 3.10: https://peps.python.org/pep-0636/
> syntax for partial application...
from functools import partial
somefunc_arg1_arg2 = partial(somefunc, arg1, arg2)
> ...and compositionA native compositional syntax would be nice.
> a typing system which can express structural types
# mypy will typecheck this code and see `MyString()` has a `.read()` method, so it's a `Readable` even though it doesn't sub-class `Readable`
from typing import Any, Protocol
class Readable(Protocol):
def read(self) -> Any: ...
def read_something(something: Readable) -> Any:
return something.read()
class MyString:
a_string: str = "something"
def read(self) -> str:
return self.a_string
read_something(MyString())
> a generalised list comprehensionI agree this would be nice.
that isnt syntax for partial appication, and the protocol system for structrural typing is syntactically and practically absurd
the question is "why do data processing libs in python look syntactically illegible" and the answer is largely, as youve posted above, that python doesnt support useful syntax
It has a rather different api, and is significantly faster. Highly recommend it.
Or you can convert it directly into a pandas Df
So you are one `df.to_numpy()/df.to_pandas()` away to `X` libary you want to use.
https://github.com/otsaloma/dataiter
Here's a comparison of dplyr vs. Dataiter vs. Pandas, which should give quick overview of the similarieties and differences.
https://dataiter.readthedocs.io/en/latest/_static/comparison...
Background: I first gained some experience with J, where I first learned to appreciate the advantages of array languages. The main advantage is actually having your own way of thinking about processing multidimensional data. Now, recently, I've gotten into the notation of APL, and it's really even cooler than the ASCII "noise" of J. The symbols make it easier for me to both write and read programs. Admittedly, for more complex operations with data it takes a lot of learning that you usually don't have. But for simple transformation APL is already quite fast usable and convinces beyond the "mainstream".
And then I read TFA and realize that I actually dislike people writing poorly formatted method chaining pandas code. The examples in the post are really nicely formatted and easy to read!
Full disclosure - the downside of modern R is the stack traces have been getting worse for a while.
"if you want to do more than statistics, let’s say deployment and reproducibility, Python is a better choice." https://www.guru99.com/r-vs-python.html
I do want to note that "there are many things for X in language Y" isn't necessarily a positive thing. It often means the community lacks the clarity of thinking or will to converge on one excellent product. Instead there are lots of okayish things each developed by a single person or handful of people.
.pipe(lambda df: (df, pdb.set_trace())[0])
I think you'll find a similar selection bias if you ask HN commenters what they think about Excel.
I think this also is a fundamental distinction between normal SWEs and People who use Excel as well as Data Engineers.
You always start with data, and you have no control over it. So this means:
1. you need state base programming env (excel, Jupyter) 2. you need to look at it to see whats there (plots)
I guess HN is mostly comprised of SWEs who built the DBs and Websites that create the data that then gets consumed by the Data Engs and the Excel people :D
If you're an analyst who knows some VBA, that can be super useful but it would probably be a mistake to try to make your VBA-driven applications bullet proof. Nobody wants you to spend that much time on it, and the odds that something completely out of your control will change and break it anyway are quite high.
Not sure it's because the article is relatively old (2016).
Not to mention, it's a lot more debuggable this way (which is generally the biggest downside to most specialised chaining approaches).
Perhaps the successor should be contemporary Pandas, or postmodern Pandas. :)
( df.pipe(went_up, 'hill')
.pipe(fetch, 'water')
.pipe(fell_down, 'jack')
.pipe(broke, 'crown')
.pipe(tumble_after, 'jill')
)is much better then something like that
df = went_up(df, 'hill')
df = fetch(df, 'water')
df = fell_down(df, 'jack')
df = broke(df, 'jack')
df = tumble_after(df, 'jill')
Would really like to hear an opinion about that.
Though if you replaced each subsequent line with df1 and df2 and so on he wouldn’t mind as much.
I can’t opine as to whether one approach or the other is intrinsically better. But echoes of his tirades still ring when I see the same variable name redefined.
In this case I would not think re-assigning df as being in violation those principles. Might even be useful if it plays better with how the code interacts with debuggers and version control.
Its not really re-used in the sense that would motivate making up new names for the intermediate steps. It’s clearly just used as a syntactic aid for the operation chaining. So in my mind both expressions are equivalent from that perspective.
I would probably insist on limiting its scope to precisely that expression though, to maintain that obviousness.
# Per store revenue for high price items
df_agg = (
df1
.query("unit_price > 400")
.groupby('store_id','store_name')
.agg({'revenue': np.sum,
'customer_id': 'nunique'})
)Groupby->agg is a good one to chain because you don’t want the intermediary almost ever.
Edit: also, I never understand people who use query. That just seems like it’s begging for problems later on. .loc for life.
How are they a replacement for pandas ? I thought the would or at least could wrap around them for scheduled execution / chaining. You would still need a Data frame handling glibrart no?
I’ll chain a few methods but never more than can easily fit on one line. Usually just something like groupby agg rename.