Lets-plot: An interactive Python plotting library using ggplot's API
github.com
github.com
Findings:
I honestly don't get the hype around R's plotting capabilities. ggplot2 is nice, but far too magical to be easily understood by beginners.
Python offers more intuitive memory management than R (which is really more a statement on how bad R is at memory efficiency).
R blows Python out of the water when it comes to expressing concise and readable linear algebra/stats computation. Done right, base R looks like mathematical pseudo-code.
R's non-standard evaluation is an under-appreciated killer feature. I've seen complex and powerful DSLs implemented implemented in base R. Python has some support for this kind of thing, but nowhere near the level of R.
I tried the Tidyverse and found it to be too magical and unstable. In general, the R package ecosystem is weird. Look at Stack Overflow and you'll find people touting all sorts of different packages to do simple operations. Needing to download a million random packages from grad students of the internet is a recipe for disaster.
http://blog.moertel.com/posts/2006-01-20-wondrous-oddities-r...
Of course, I reserve the right to change my opinion after I build a few things. Dynamic typing... hmmm...
I haven't found that to be the case. Loading a 2gb CSV file when you have 8gb of ram is touch and go with pandas but with data.table it’s a breeze not to mention operations once loaded up are a fair-bit faster. The pythong version has recently been released and already provides a serious speed bump over pandas not to mention out the box memory-mapping for huge files. I recommend checking it out if you haven't heard of it https://h2oai.github.io/db-benchmark/
In particular, I don’t see how the tidyverse is at all ‘magical’ - take the two most popular tidy libs, dplyr and ggplot2. The api for dplyr is very explicit, intuitive, and based on long-standing precedence (sql). The api for ggplot2 is admittedly less intuitive, but is itself an implementation of widely known framework (Grammar of Graphics). If by magical you mean in their abstractions, those are about as far from magical as one can get.
The tidyverse does get iffy sometimes when you need to dig into the rlang/tidy eval area, but that’s all well-documented.
I haven’t had to use R in a production environment for about a year now, but I always enjoyed the concise tidy api.
Base R is a total mess and largely inconsistent, but I guess that’s what you get when statisticians from different uni’s patchwork a language in their spare time.
(I know some people swear by data.table, but I find it gets ugly fast if you want to do anything more than simple joins or split-apply-combine)
Non-standard evaluation, in my opinion, does fine in the top layer (analysis scripts, DSLs). But it's a pain to build on top of magic without a more concrete layer between. Base R uses character vectors and name attributes to hide the magic. The Tidyverse uses symbols. I find character vectors much easier to understand and manipulate.
Formulas are the only scenarios in base R where symbol and expression manipulation is sometimes necessary. And I don't do that thinking, "Boy, I wish this was how I did 95% of my work."
Personally, I use data.table. It's been great even for packages. If it's an internal package dataset, I don't need to manipulate symbols. And if it's a user-supplied argument, there's usually a simple and fast way to do it with a character vector. Or I define an S3 class, give a helper function that produces a dataset with known names, and just write functions around that. The user can handle combining data however they want.
>Base R is a total mess and largely inconsistent
I wouldn't go so far as to say "total mess." Not even an annoyingly big mess. The few functions that make up most of analysis or package code are dead simple. Inconsistent argument names is hardly a problem with an IDE.
This is baloney. You don't need to rely on a million R packages. You really only need a good plotting library (the built-in one is fine, or use ggplot), a good data frame library (dplyr, data.table, or use build-in one), and a few extras to handle some weird rough edges (lubridate, forcats).
This takes you to 90% to your analysis goals.
Only novices are 'using a million random packages', which is probably the case for python as well.
Though to be complete, any ML project means python and scikit learn.
Is there any coincidence that there are literally tens of different (and concurrent) rendering libraries for Python while the R world more or less settled for ggplot2? Writing matplotlib is such a pain in the ass that people go to great lengths to actually avoid using matplotlib (while still not having the expressivity, features and ease of use of ggplot2).
Magic is bad when you're shipping production code that's shared among multiple people who have to then spend a lot of time assimilating the mental model of the magic. It's perfect when you just want to plot stuff and draw nice figures.
[1] http://karolis.koncevicius.lt/posts/r_base_plotting_without_...
Are you including Numpy as part of Python in this statement? I've found Numpy to be as mathematically expressive as R, although I'm not a R power-user.
I wish more projects would use that way of thinking but it seems that in the jupyter/julia/python world there's too many choices and they all attempt a "kitchen-sink" approach for visualization.
It’s interesting that you say this. I’ve been using R daily for almost 3 years now, and I was originally taught the Tidyverse.
However, I also tutored in biostats and have collaborated with many different faculty and students.
On one hand the Tidyverse created an opening for R learners especially, but it has lead to some controversy as well.
Trust me when I say there are plenty of R users who have never heard of or used the Tidyverse.
It leads to some difficulty in teaching/collaborating because do I use grepl or str_detect? lapply or map? Do they know the magrittr syntax (%>%)?
Then there’s non-standard evaluation, which often forces some arbitrary meta-programming on novice programmers.
R is foremost the stats language, and tidyverse has attempted to make it more general purpose, for better or worse depending on who you ask.
True story: writing tidyverse functions is so hard that I once worked for a company with business critical code running in incredibly long R-scripts with almost no functions, and 100 line pipes.
While that may be an extreme example, base-R is much, much easier to get started writing functions for (like I've been using R for a decade now, and the "idiomatic" way to use NSE with ggplot and dplyr has changed multiple times over that time).
Meanwhile, base-R is fugly, but its solid and backwards-compatible to a fault.
However, once they get that right, then tidyverse is close to perfection
I think the newish “interpolation” syntax with double braces {{ }} is getting there.
Damnit Hadley, why must you do this to me?
(Really I'm bitter because I know that I'll end up maintaining a bunch of important code in version N-2 of Hadley's NSE adventures at some point in the future).
https://www.tidyverse.org/blog/2020/02/glue-strings-and-tidy...
There's a huge shadow universe of scientists running particular R packages for these analyses, which never gets seen from HN/programmers.
While Python is definitely (sadly, IMO) a better language for building ML systems (because more people know it, and it's harder to write impenetrable code), R is definitely a better DSL for statistical analysis, modelling and graphics.
It's a shame that R takes its lineage from a language developed in the 70's (S), as that's a cause of many of the inconsistencies in the core language.
I suspect, however, that it would not be anywhere near the top 10 of the TIOBE index like it is now and that it's userbase would consist of mostly of statistics practitioners rather than huge swathes of people that need to perform any basic data analysis.
Also for a lot of people R is going to be the first and often only language they're going to learn because they have no use for programming outside of statistical analyses and it's damn practical for that single use. Right now generations of students are being taught in R, right now bioconductor is booming and being added to everyday. That stupid arrow assignment operator has plenty of good years ahead of it.
here's the path to Altair scrapping 20k lines of code to make the api simpler - https://twitter.com/jakevdp/status/1006929128119926786
Altair is really good. Probably as good as ggplot2
I can think of 3-4 off the top of my head.
Would be curious to hear which ones you're thinking of (since I wouldn't be surprised if the are others, but also am guessing no others hit both points)
Plotnine is the only one another human has told me about, though.
My (wildly speculative) guess for the motivations behind this library would be that Kotlin, as a base, is usable in the other JetBrains IDE's, e.g. reusing the same library and grammar when developing in Java or Ruby.
> The Lets-Plot for Python library includes a native backend and a Python API, which was mostly based on the ggplot2 package well-known to data scientists who use R.
> R ggplot2 has extensive documentation and a multitude of examples and therefore is an excellent resource for those who want to learn the grammar of graphics.
> Note that the Python API being very similar yet is different in detail from R. Although we have not implemented the entire ggplot2 API in our Python package, we have added a few new features to our Python API.