ggplot2 is obviously fantastic and makes beautiful plots, and very easily at that. However it is definitely a "convention over configuration" tool. For 99% of the typical plot you might want to create, ggplot is going to be easier and look nicer.
However matplotlib lib really shines when you want to make very custom plots. If you have a plot in your mind that you want to see on paper, matplotlib will be the better tool for helping you create exactly what you are looking for.
For certain projects I've done, where I want to do a bunch of non-standard visualizations, especially ones that tend to be fairly dense, I prefer matplotlib. For day to day analytics ggplot2 is so much better it's ridiculous. The real issue is that Python doesn't really offer anything in the same league as ggplot2 for "convention over configuration" type plotting.
Fully agree on Pandas. R's native data frame + tidyverse is world's easier. Pandas' overly complex indexing system is a persistent source of annoyance no matter how much I use that library.
Is it just the syntax/readability that annoys you, or are there actually problems that need like n steps more to do the same with Pandas?
If you're mostly dealing with Neural Nets you won't see much R, but for anything really statistical in nature R is a much better tool than Python. For anything that ends up in a report R is much better than Python (a lot of very valuable data science work ends up being a report to non-technical people).
> breaks down on data manipulation
This is very outdated. The tidyverse eco-system has bumped R back into being first in class for data manipulation now. This becomes less true as you get further and further from having your data in a matrix/df (I can't imagine doing Spark queries in R), but if you already have a basic data frame, manipulation from there is very easy.
Even for things that end up in production, whether you're in R or Python, whatever your first pass is should always be a prototype and will have to be reworked before you get close to moving it to production.
Depends on your definition. While not very often 'deployed' in 'production'. I know lots places in all kinds of industries where people reach for R as soon as they have to look at some new data.
You are right in the sense that R is typically not used end-to-end as far as I can tell, but already tries to start with a data connection to some sort of dump or datalake, or datawarehouse.
Many people in my team use Python for modelling, but grab ggplot in whatever way to make their presentations and visuals (they all use different methods, usually something messy like mixing python and R in a notebook or so). GGPlots also has a vast library of super high quality plugins.
Python is far far behind in the viz space
I primarily use Altair within Python. ggplot is ahead of Altair in some respects, but behind in others.
For example, here is a chart which can be made in Altair:
https://altair-viz.github.io/gallery/seattle_weather_interac...
Note:
- You can brush over the date range to filter the bar chart
- You can click on weather type to filter the scatter chart
- It can be embedded in any webpage with these interactive elements in tact. Since the chart is represented by json and rendered by javascript, the spec also embeds the data within the chart itself, and allows the user to therefore change the chart however they want
You can even build something like gapminder: https://vega.github.io/vega-lite/examples/interactive_global...
More examples here: https://altair-viz.github.io/gallery/
Matplotlib, though ... that's a harder beast to internalize. I know it's possible to make high-quality matplotlib plots, but it's much harder for me. Like pandas, it's a library that I don't want to denigrate because I know people put lots of effort into it, but I can't lie -- I'm not a fan.
That said having used both the DSL's for plotting and data wrangling in the R package ecosystem are vastly superior to pandas and python plotting libraries. For modeling I actually like the better namespacing of Python which helps keep things more legible when there are a ton of model options to choose from, assuming you don't need cutting edge statistics.
It's not that much harder. There's no pytest, but testthat works well enough. I've developed a few packages internally in R and wouldn't say it was that much harder to ensure correctness than for the corresponding Python packages. (We used to keep them in sync, before basically moving everything to Python.)
You also have the dump.frames option, which will save your workspace on failure, which is incredibly useful when running R stuff remotely/in a distributed fashion.
Both in terms of conciseness and performance
It depends, I've worked in some places where R was the core part of their data infrastructure. Data manipulation (of non text) is far, far better in R.
Integrating with other systems can be tricky though, and you don't have the wide variety of Python libraries available for core SE tasks, so it can often make sense to use Python even though it's not as good for a lot of the core work.
Additionally, R is a very, very flexible language (like Python), but without strong community lead norms (unlike Python) so it's pretty easy to make a mess with it.
Finally, when you need to hand over stuff to software engineers, they vastly tend to prefer Python, so it often ends up being used to make this stuff easier.
Like, in R there's a core tool called broom which will pull out the important features of a model and make it really easy to examine them with your data. There's nothing comparable in Python, and I miss it so so much when I use Python.
That being said, working with strings is much much nicer in Python, and pytest is the bomb, so there's tradeoffs everywhere.
> Additionally, R is a very, very flexible language (like Python)
I'd argue that R is much more flexible than Python syntactically. There's a reason that every attempt at recreating dplyr in Python ends in a bit of a mess (IMO) -- Python just doesn't allow the sort of metaprogramming you'd require for a really nice port. Something as simple as a general pipe operator can't be defined in Python, to say nothing of how dplyr scopes column names within verbs.
Arguably this does allow you to go crazy in a way that ends up being detrimental to readability, but I'd say overall it's a net benefit to R over Python. I really miss this stuff and have spent an undue amount of time thinking of the best way to emulate it (only to come up with ideas that just disappoint).
> Finally, when you need to hand over stuff to software engineers, they vastly tend to prefer Python
Indeed, this is maybe 50% of the reason my organization has pushed R to the sidelines over the past few years. We used to be very heavily into R but now it has "you can use it, but don't expect support" status.
> Well that's just lazy evaluation of function arguments, which can't be done in Python.
"Just lazy evaluation"! :) It's a pretty big deal. This is three-fifths of the way to a macro system.
> But if take a look at the Python data model, it does seem super, super flexible.
Sure, you can have a lot of control over the behavior of Python objects (some techniques of which remain obscure to me even after using Python for many years). But you don't have anything like syntactic macros. You can define a pipe operator with macropy, though -- it's pretty easy. But macropy is basically dead now I think (and a total hack).
> You'll still need strings for column names in any dplyr port though, because of the function argument issue.
This is major, though, because you can't do this:
mutate(df, x="y" + "z")
You have to do something like what dfply does, defining an object that defines addition, subtraction, etc. mutate(df, x=X.y + X.z)
But that hits corner cases quickly. What if you want to call a regular Python function that expects numeric arguments? This won't work: mutate(df, x=f(X.y))
etc. Granted, this only really works in R because it's easy to define functions that accept and return vectors. So in that sense it's kind of a leaky abstraction. But you couldn't even get that far in Python, because X.y isn't a vector ... it's a kind of promise to substitute a vector.Give Python macros, I say! To hell with the consequences!
Not yet, but there’s a PEP for that:
(Why can I reply at this level of nesting now, whereas before I couldn't?)
Fundamentally though, both DS Python and R are abstractions over well-tested Fortran linear algebra routines (I'm sortof kidding, but only sortof).
Well that's just lazy evaluation of function arguments, which can't be done in Python. But if take a look at the Python data model, it does seem super, super flexible. You'll still need strings for column names in any dplyr port though, because of the function argument issue.
Like, both Python/R derive from the CLOS approach (Art of the Metaobject Protocol), but R retains a lot more of the lispy goodness (but Python's implementation is easier to use).
Then when doing basic operations like "group by" you end up excessively elaborate indexes that are in my experience useless and always need to be manually squashed to something coherent.
It's a common joke for me that whenever even a seasoned Pandas user cries out "gaarrr! why isn't this working!?" I just reply "have you tried reset_index?"... this works in a frighteningly large number of cases.
Generally, I've even preferred Spark to pandas, though it's hardly less verbose. Coming from R, it's much slower than data.table and nowhere near as slick and discoverable as dplyr. Its system of indices is a pain that I'd rather not deal with at all (and, indeed, I can't think of another data frame library that relies on them). I hate finding CSVs that other data scientists have created from pandas, because they invariably include the index ...
Handles time series really well, though.
Recently I've been using polars (https://github.com/pola-rs/polars). As an API I much, much prefer it to pandas, and it's a lot faster. Comes at the cost of not using numpy under the hood, so you can't just toss a polars data frame into a sklearn model.
That being said: > I hate finding CSVs that other data scientists have created from pandas, because they invariably include the index ...
This is also default in R, with row numbers (like I have ever needed them). To be fair, it's gotten better since people stopped putting important information in rownames.
Polars looks interesting, thanks for the recommendation!
Ideally you should be using the parquet format which will use the binary format, preserve column types and indexes [df.to_parquet(<file>); df = pd.read_parquet(<file>)]
You can get away from a lot of problems by simply avoiding text files
More generally, the API is large, all-consuming and not consistent. sklearn is best in class here, I rarely need to look things up whereas the pandas docs autocomplete in my browser after one or two characters.
I like to stick to basic, widely used tools when possible so I'm biased against it versus just wrangling it out with matplotlib. But proplot does look compelling, like it was written for exactly my complaints.
I will say that my very first "real" programming experience was Matlab at a research internship, so maybe i just got used to working in vectors and arrays for computational tasks.
Have you worked with R? R, like matlab, natively supports vector based operations. In fact, all values in R are vectors. Many of the problems with Pandas ultimately boil down to the fact that you have to replicate this experience without truly being in a vector based language.
Pandas indexing makes sense once you get it, but it does seem to require a lot more words than equivalent statements in R.
My primary language is python, but I have been picking up some R.
I would be hard-pressed to find a working data scientist whose definition of data science is "that thing you do with sklearn, Deep Learning and Numpy".
Most problems are mostly tabular, IME.
I completely agree that text, images and video are much, much better handled by Python (that's why I use and know both).
It could be. It's such a broad job title and it looks so different across different companies and teams that the main tool for one data scientist might be something that another data scientist never has to touch. Different data science jobs prioritise different tools, that's all.
Still, if anyone here has managed to find a data science job in which tabular data management is not a sizable piece of what you do, I'd like to know some details!
Still, my suspicion -- at least from my corner of data science -- is that such individuals are rare, and that most data scientists do make use of tabular data more often than not.
I'll also add my vote for the superiority of data.table and ggplot2 to any Python alternatives. the bloat and verbosity of pandas is a daily struggle
Yessss. I loathe indices, and have never been in a situation where I was better off with them than without them.
> Regarding fast you have something like Vaex on python sid
I've never used Vaex, but I've used datatable (https://github.com/h2oai/datatable) and polars (https://github.com/pola-rs/polars). Polars is my favorite API, but datatable was faster at reading data (Polars was faster in execution). I'll have to give Vaex a try at some point.
If you already think pandas is slow I think you'll be surprised how much more strongly you feel after using data.table!
Not really, tbh. Most of my jobs (even when the primary output was models) require spending a _lot_ of time data wrangling and plotting. R is much, much better for this kind of exploratory work.
But if I need to integrate with bigger systems (as I normally do), there's a stronger push for Python to reduce complexity and make it easier for SE's to understand and maintain (some of) the code.