What would it take to recreate dplyr in Python? (2020)
mchow.com
mchow.com
Was this just a high level (possibly misguided) paradigm that the pandas devs fell in love with - or is there a good, performance related reason to embed it so deeply in the API?
i used to write a lot of functions which went like this:
def convert_price(price_frame:pd.DataFrame) -> pd.DataFrame:
price_frame["new_price"] = np.exp(price_frame)
return price_frame
and i came back to these functions, and i was always like. hmm what needs to be in the price frame. which columns etc. Also it mutates the state of the data framewhile i now write functions like
def convert_price(price:pd.Series) -> pd.Series:
return np.exp(price)
However what is good as well you can treat the series as a single dimensional array and do operations on it. Its not a perfect example since im not using the ids haha but you might see what i mean :DI am not sure if I am just older now but I am more and more set in my dplyr ways and it’s hard for me to adopt the python way of wrangling data frames.
I imagine this is mirroring R's `write.csv()` behaviour.
But I agree, if you're designing something sane you probably shouldn't copy R.
I believe this was because Pandas' initial primary use case was manipulating time series (datetime-indexed numerical vectors), which are used extensively in financial institutions such as hedge funds and trading firms (Pandas was initiated as a skunkwork project in AQR Capoital Management). You can see the lineage in its very extensive collection of convenience methods for manipulating time series (pd.Series where the index is some datetime type). A pd.Series is a numpy array with a meaningful index, and a pd.DataFrame is a collection of pd.Series with a shared index. If you use Dataframe to store and manipulate multivariate time series, the api is quite sensible. So pd.Series and pd.DataFrame are probably the datatype that you'd design to store time series of (a portfolio of) stock returns. (Old foggies who'd use something like Matlab before know that just being sure your vector/matrix calculations are using correctly aligned dates was not a given.)
Dplyr and its R data.frame heritage are what statisticians would probably use to record measurements / experimental outcomes on individuals. There is usually no meaningful index/natural primary key, and the order usually doesn't matter. It's much closer to a relational database table (unordered collection of tuples), but for analytics rather than transactions so column- rather than row-oriented.
For data tables without a meaningful natural index, the Pandas api is much more confusing and cumbersome than needed. It happens that a lot of ML applications fall in that category, but during the early 2010s when Panndas took of it had very little competition.
https://www.dlr.de/sc/Portaldata/15/Resources/dokumente/pyhp...
When you start with simple data frames it feels dict-like and very pythonic —- zero learning curve. But then multi-indexes are lists of tuples and things get tricky. It’s also important that data frames always have an index whether you realize it or not, the default being a RangeIndex.
This is why the “to_csv” bites people, there’s a numerical index. Once an index is set you no longer have to say “index=False”.
With the default RangeIndex “df.loc[0]” and “df.iloc[0]” will give the same result because the first record has both the position 0 and index value of 0. Once a “meaningful” index is set you need to start referring to it by value rather than position. This makes it much easier to manipulate data, for me at least.
Not least in that both broadcasting and pandas indicies seem surprising and magical if one doesn’t understand the details.
I just wish Pandas treated named index columns as columns. When I write df["blah"], I don't want to have to remember whether I just loaded the data and it's a normal column or if I just grouped on "blah" and it's an index column.
Currently, in the latter case, subscripting doesn't work and you either have to do a reset_index() or look up the correct incantation -- something like df.index.get_level("blah").
I have a feeling that the root of the problem is that Pandas ended up the same concept of "index" both for optimized lookup and for the UI of grouping / joining. My guess is that get_level is less efficient than it should be, and thus Pandas discourages using it by making it obscure.
Ditto the rest of the tidyverse.
Comparison to dplyr, ...: https://dataframes.juliadata.org/stable/man/comparisons/#Com...
Comparison to Pandas: https://dataframes.juliadata.org/stable/man/comparisons/#Com...
using DataFramesMeta, Statistics
@chain mtcars begin
@select :Cyl :HP
groupby(:Cyl)
@transform _ begin
:dumb_result = myarbitraryfunc.(:HP)
:demeaned = :HP .- mean(:HP)
end
end
(mtcars is from RDataSets.jl)[1] https://juliadata.github.io/DataFramesMeta.jl/stable/dplyr/
``` def sqldf(df: DataFrame, query: str) -> DataFrame: ... ```
Glad to to see duckDB delivering, finally, on the promise of running SQL against in-memory dataframes
If its not zero copy. It is still not a big deal. Pandas make a lot more copies internally. I truly wouldn't worry about that single copy if you have a order of magnitude speedup overall.
That said, it’s so much better than pandas for data manip that I’ll probably still try to use it.
Are you the author? If so, thanks for being so responsive on GitHub. You fixed basically every issue I had almost immediately back when I was learning Polars. It was awesome.
But I will improve it. ;)
Does your setup allow for an end-to-end solution? I mean, can I sink time into that setup and feel like I have everything I need to for regular data-wrangling?
I'm sure Pandas is amazing, but as a newbie I found myself doing many transformation logic with python data structures because it's just so much easier.
Maybe I'm dumb but going around the docs sometimes was like :/
(I haven't really used it, but it looks promising)
Maybe someday Python'll get a macro system ...
Siuba has come a long way since I wrote this, and now can optimize for fast grouped operations!:
* https://github.com/machow/siuba
* https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
As other have said, escaping pandas is hard. Many visualization and data manipulation, validation and analysis libraries expect pandas input.
Siuba is really cool in that it offers a convenient syntax on top of pandas (and SQL databases) without requiring its own data format.
out_rec = []
for id, group in data_frame.groupby("id"):
ladidida....
result = f(group)
out_rec.append(result)
in my experience it isn't much slower than a groupby.apply.* If f() converts grouped data to something ungrouped, then you can't use a similar function f2(f(group))
* If f() returns a grouped object, then you can't do basic operations like f(grouped) + 1, because DataFrameGroupBy, SeriesGroupBy do not define basic operators like addition. Let alone operations against other grouped data.
A lot of this is worked out in siuba now, and this doc explains a bit more:
https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
(user_courses
.set_index(["student_id",
"course_id"])
.unstack()
.apply(lambda x: x+1))(no disrespect to to the package in the article or OP who I know is active in this thread. just a general motif that I keep coming across in python).
There are a lot of problems you encounter when using arrays for data analysis, like some of their funky behavior with strings [0], but it seems like extending arrays, or building new types of numpy arrays would have been better than new data structures like the pandas Series.
(pandas folks thought a lot about these problems so I could be very wrong).
Port the functionality of the R package but try to keep it python. Run flake8.