Self-Directed Pandas Crash Course
kellyfoulk.herokuapp.com
kellyfoulk.herokuapp.com
In particular, I find this answer infuriating [1]. I've come across it so many times. Look, I have a CSV with 200 rows and I need to loop through them in the most intuitive way. Sure, its not optimal but I don't want fast code. I have a mental model of how to modify this dataframe. Let me do it, please.
1. Make a good old numeric for loop like "for ii in range(len(df))" with df.iloc or df.ix, etc.
2. Use df.apply() to create a new DataFrame with your changes.
Both of these are mentioned in brief answers to that SO post. But not in the accepted answer. A lot of the answers focus on the most efficient ways to do things, even though the question was very basic and did not imply the data were large.
That being said, Pandas isn't a very good array language. I really like kdb+/q and find it better and more expressive than almost any other language I've used.
Pandas to me are still big lazy black and white bears eating bamboo, until someone can point me to something more intelligible.
That said, one of its best features, and probably the only thing I use it for these days is, `pd.read_sql(sqlstr, conn).to_csv(fp)`. This is far less cumbersome than using psycopg2.
Edit: for charting these days, similar to OP's visualiztions, I highly recommend vega-lite.
If you really just want to use loops and stuff (which I would discourage) just use a list or a dict or something.
Like my post said, I favor vanilla python over pandas... so yes, I use lists, dicts and somethings. FWIW, though, my workflow pushes everything into postgres, and things that would normally go into pandas are just accessed through SQL through and with helper functions.
It’s ‘read_<format>’ to read something in, but ‘to_<format>’ to write it out? In which world is this intuitive? Surely it should be read/write or even from/to?
Imagine you were using Pandas for the first time, working through a tutorial, and you've just learned that pd.read_csv reads a csv into a dataframe.
Intuitively, what would you expect the corresponding output function to be called? I'd hazard that a vast majority of people would guess at some version of write_csv, and would experiment with either pd.write_csv, or df.write_csv.
Pandas is really a DSL unto itself and is heavily influenced by R, where the same dynamic happens. Programmers coming from a background where procedural control flow constructs are basically second nature bump up against statisticians for whom array-based programming (in the form of overloaded mathematical notation acting on both scalar and vector values) is second nature.
R and pandas are both very array-oriented programming languages (the most extreme example of this might be early-era APL) and it's really going against the grain to implement things with explicit iteration.
It's kind of like trying to program in Python without using loops or list comprehensions and asking just how to do everything in recursion. You can... but someone is bound to point out that doing everything with recursion (and the concomitant trampolines to prevent stack overflows) is not the Pythonic way.
(Also separately @dang, I feel like I'm running into a very minor bug with time stamps, where when I'm composing this reply I get "9 hours ago" for systemvoltage, but in the main thread I get "3 hours ago")
About `pd.read_sql()`. That's totally awesome.
I realized that Pandas is a very useful tool for many thousands of developers, I only have a problem with its interface. Obviously, I've been using it for many years for a reason!
You are operating on n-dimensional arrays, not elements, so your need to write code that expresses that intent.
Anyone who thinks that anything besides writing "a + b" to add two matrices together is a good or simple solution is crazy.
Operate at a higher conceptual level. Don't use loops. Transform and compose your data, not your datums.
It's a balance, and I think pandas is on the more complex side of things.
Those might be more excuses than reasons, but it’s my experience.
My problem is that this is super inconsistent. Some things are done as a method call on an object, others by passing the object to a pandas function and others yet by passing a function to a method on an object. This is the major source of frustration for me.
Maybe there is some logic to that, but I haven’t found it yet and I think that is a sign of bad API design. Its like PHP to me. All nice and documented but useless without Googling everything
I've also recently found that using the sqlite3 command line tool increases my productivity when doing data sciency stuff a lot. It's a super fast, super simple way of making sense of CSV data, especially for the selecting an joining operations that are so unintuitive in pandas. Once that's done, I can dump the results into Jupyter or RStudio for transformation and/or visualization.
Another personal productivity win I've discovered is using two-line python scripts to write really long and repetitive SQL commands (select count a, b, c... from huge_table where d... and e... and f...; etc.) and then running them in SQLite.
It's not like I'll remember all the syntax, but good to know what tools/tricks exist.
1. The style of your website makes it pretty much impossible to tell if a bit of text is linked anywhere. I only figured it out after clicking on the word "here" in the first paragraph and coming back and clicking on essentially everything. None of this happened until after I visited the site on my desktop instead of my mobile and slowed down to read things carefully.
2. It would be great to get a link to the denvergov.org data set, or corresponding area of the site.
One shouldn't have to, but FWIW the HTML of his page is very clean so you can see all of the links by viewing the source.
I saw it on: https://www.reddit.com/r/Python/comments/lain0r/hey_reddit_h...
* Self-Directed, Pandas Crash Course
* Self-Directed Pandas, Crash Course
Rewriting as Self-Directed Crash Course on Pandas would eliminate the ambiguity.[1] https://youtube.com/playlist?list=PL-osiE80TeTsWmV9i9c58mdDC...