Effectively Using Matplotlib (2017)
pbpython.com
pbpython.com
Pandas also comes off as an unintuitive joke, but my displeasure with it has mostly worn off. Matplotlib however makes me feel angry pretty much everyday.
Usually this is resolved by a complete redesign of the software that gets delayed forever, never becomes production ready and eventually disappears into obscurity.
Tidy verse assumes tidy data. If you are not working with tidy data, it is unlikely to be a big help. Most data can probably be thought of as tidy.
Remember that any and every operation on a data frame returns a data frame, so unlike chaining in Pandas, you never have to worry if a method you want to use belongs to a series or a data frame, or if your method is returning a series or a data frame.
Select() selects columns, filter() selects rows. This never changes unlike the [] which means different thing depending on if it is used on a data frame (which you are not guaranteed to be served after calling a method on data frame in pandas!) on a series or using the .loc or .iloc methods.
There is no index, instead you just filter on rows.
Pandas comes with a ton of build in utilities which the tidyverse doesn’t, mostly because R is already full of functions you can easily apply across columns.
But particularly pandas date handling functions are really cool
My absolute favorite DataFrame library is saddle (for Scala), which I helped write at my old quant job. Very FP oriented and an absolute pleasure to use. Though maybe it’s no surprise that I like something I worked on.
An incomplete list of things that I dislike about Pandas are:
Too many parameters and knobs for each function
Inconsistenty between inplace and copying operations
Unintuitive function names compared to FP
Too much magic in how things work
Functions and parameters accept a wide range of types in order to make things “just work.”
Lots of non-orthogonal convenience functions that do mostly the same thing
I’m not familiar with how “normal” Python is written, but I suspect a lot of the problems come from the abuse of dynamic typing. Dynamic typing allows you to just add more and more levels of crap without actually changing your data/type model. I think there’s a lot of value in “correct” APIs, vs convenient ones.
That being said, Pandas is extremely powerful, and usually very succinct. Maybe not as nice as kdb+/q (nothing really compares for time-series data), but still pretty good.
Honestly, for 99% of uses Seaborn is great, so long as you remember to use the latest version---for some reason, a lot of people seem to have 0.8.0 installed, and the api changed with 0.9.0.
For uses beyond what Seaborn can do, I think that the best strategy is just to figure out a personal plotting language and then wrap that up into a personal library so you never have to think about that again. That's kind what I've done: I threw together a library to produce some basic figures that are suitable for printing,[1] and now I never have to think about those figures again.
If you can get away with it use pandas' plot, seaborn, altair, etc.
You could always even just write out GNU Plot commands and then call it. I used that for the test harness I wrote in the robotics club (and everything was in C!) to plot the trajectory of the robot in auto mode. It’s super easy! I don’t remember if the GUI has all the panning and scaling though.
Since their 4.0 release (https://medium.com/plotly/plotly-py-4-0-is-here-offline-only...), there's no longer any connectivity to their cloud service, it's "offline only". It used to have an offline interface _and_ a connected interface, now it's offline only.
See https://plot.ly/python/is-plotly-free/ for full details.
Every single time! Unlike, say, numpy, where everything is consistent, makes sense, and works as expected almost always.
What I can't decide is if: a) matplotlib is difficult; b) plotting is inherently complex like writing sheet music; c) object oriented programming leads to gratuitous complexity. So I chalk it up to some combination of the three, but have never felt compelled to try anything different.
I don't do anything for publication, but I use plotting inside of software that I use for running lab experiments, prototypes of measurement hardware, and even in the factory. So I'm using Python for what people would have used LabVIEW for in the past. My programs need to produce readable plots without tweaking, because I don't know in advance what the data are going to look like. The combination of tkinter and matplotlib is really huge for me.
I just went through this process myself (http://kachess.k2company.com) and this outline would have been SO helpful. I learned these points the slow, hard way.
(And while it doesn't suck, it's certainly not fun and intuitive.)
Matlab plotting is extremely powerful and versatile. Sometimes the output could be nicer but the interactive figure hierarchy is great. Matplotlib on the other hand is, at least to me, a lot more clunky to work with. But it gets the job done and the output often looks nicer and solves my gripes with Matlab.
In my mind, ggplot2 is actually more comparable to matplotlib in that it's more expressive, but less intuitive. It's interesting that the 'successor' to ggplot2,ggvis (which is on ice) , used Vega as a backend.
I just wish that its documentation examples would consistently provide the OO interface version of how to achieve each example, at least alongside the state-machine version.
It's always frustrating to see an example image that shows exactly what I want to achieve, and then click on the code for it and it's using the other interface, and I have to try to guess the equivalent OO commands. Which are always slightly different, like set_ylabel instead of ylabel...
Specifically seaborn’s catplot (for categorical), lmplot, swarmplot, pairgrid, and facetgrid.
The seaborne gallery really has an extra level of expressiveness that you might not have considered as an amateur visualizer and you can make some very nice things.
Matplot lib runs underneath it so you’ll need to learn all of the adjustment functions: lim, figsize, ticks, etc. but I think it’s fine overall.
Charts are hard because there’s more depth than people realize and if the library wasn’t deep you’d be unable to express that depth.