If this also turns out to be inscrutable I may be forced to conclude that I'm stupid...
If this also turns out to be inscrutable I may be forced to conclude that I'm stupid...
You need to work with pandas consistently for a month or two, and then it'll all click.
pandas is not complex, nor deep. It is, however, very broad. Most of the time it is "Here's what I need to do. I'm sure there's an API or two in pandas that will let me do this," and then you spend an hour or so looking at the documentation to find those APIs.
My first month or two was: "I need to do this. Let me Google". Pretty much every time someone had asked that same question on SO.
If you stick to it for 2 months, you'll eventually "learn" all the routine tasks and Googling stuff becomes only occasional.
And it does help if you're familiar with NumPy.
Modern pandas is a bit more idiomatic now though: https://tomaugspurger.github.io/modern-1.html
Pandas is basically an R data frame for Python. A sloppy description of that is a text mode spreadsheet.
The description of Bonobo doesn't immediately invite the comparison to Pandas, to me anyway.
Mostly, when I want a quasi-mathematical look over a dataset, pandas is my tool of choice. For all those data pipeline things that reasonably fit on one computer, I do use bonobo.
Can anybody comment how Bonobo compares to Luigi?
Is pandas the wrong kind of tool for this type of thing? Going off what rdorgueil has said, I'm beginning to suspect so. Is there a data-wrangling 'gold standard' library for python?
Create a object/class called
AuctionResult
- some datetime
- value
Then you'd query it qs = AuctionResult.objects.all()then you load it into a pandas dataframe:
df = read_frame(qs)
After that you can do all sorts of the fun stuff I imagine.
As an example from the pandas docs [1], in dplyr you can do
> gdf <- group_by(df, col1)
> summarise(gdf, avg=mean(col1))
In pandas this is similar to
> df.groupby('col1').agg({'col1': 'mean'})
But dplyr's summarize it's much more flexible than agg, as you can do all kinds of things to any number of columns. E.g.
> summarise(gdf, some_name = f1(col1) + f2(col2))
But in pandas you can apply 1 function to 1 column with agg.
[1] http://pandas.pydata.org/pandas-docs/stable/comparison_with_...
gdf = df.groupby('col1').agg({'col2': np.mean, 'col3':np.std,'col4': lamba x: np.mean((x) / np.std(x))})
once you've got your grouped dataframe, go nuts
gdf['some_name'] = gdf['col1'].apply(f1) + gdf['col2'].apply(f2)
(I should have used a better example, like summarise(gdf, some_col = f(cola / colb))
https://courses.edx.org/courses/course-v1:Microsoft+DAT208x+...
Their is some pratical exercises that you do in your browser that really helps to get the grasp of it.
Don't miss Pandas, they are really cool!
http://nikgrozev.com/2015/07/01/reshaping-in-pandas-pivot-pi...
(I´m not affiliated to site)