An opinionated view of the Tidyverse “dialect” of the R language
github.com
github.com
After doing a few personal projects using base R with complex manipulations, I would have quit using R entirely if it were not for tidyverse/dplyr.
I'm a pragmatic software engineer: I prefer to get a data analysis job as fast as necessary, and IMO there isn't much merit in making things more complicated by using base R just for the sake of base R. It's definitely not faster, both in cognitive load and in speed. In my data science work there hasn't been a single case where I've needed to fall back to base R, or wanted to.
Pipes are being undersold here: the examples only show a single piped action, whereas pipes shine when you have to do 5+ operations in a single manipulation (I've seen base R code that does that by just having a lot of redundant df assignments and it's a pain to read).
either that or a profusion of variables going into your global environment, which is a massive pain to track!
In my experience if there's a package that's required in a script/R Notebook outside of tidyverse, I just import the entire library at the start of the file Python-style (to make the dependency obvious), which would avoid this issue.
the_data <- read.csv('/path/to/data/file.csv')
the_data <- subset(the_data, variable_a > x)
the_data <- transform(the_data, variable_c = variable_a/variable_b)
the_data <- head(the_data, 100)
is a pain to read compared to the_data <-
read.csv('/path/to/data/file.csv') %>%
subset(variable_a > x) %>%
transform(variable_c = variable_a/variable_b) %>%
head(100)
I thinks it’s fair to mention how by making things more complicated by using pipes you expose yourself to other issues.(In my opinion the first variant is not less readable, and it has the advantage of allowing any of the operations to be removed by commenting the corresponding line, while in the second case to remove the last operation you also need to remove the piping operator in the preceding line.)
[Edit: I’m not sure I understand your “out of scope” remark. You don’t mean that magrittr is outside of tidyverse, do you?]
I'm a bit suspicious about that coincidence - after I met the pipe operator, a lot of my code looks like that. Before I met the pipe operator, not very much of my code looked like that.
As long as people are thinking in pipes then they can choose if they value making commenting code cheap vs making reading/typing code cheap. But it is quite important that beginners are introduced to pipes as a concept as they learn R. The semantics are critical, using the syntax is optional.
the_data <- head( transform( subset( read.csv('/path/to/data/file.csv'),
variable_a > x),
variable_c = variable_a/variable_b),
100)
All of them are equivalent, each representation has its advantages and disadvantages. the_data <- transform( subset( rbind( read.csv('/path/to/data/file1.csv'),
read.csv('/path/to/data/file2.csv')),
variable_a > x),
variable_c = variable_a/variable_b) the_data <- read.csv('/path/to/data/file1.csv') %>%
rbind(read.csv('/path/to/data/file2.csv')) %>%
filter(variable_a > x) %>%
mutate(variable_c = variable_a/variable_b)
You've thrown out a few variations on a theme here, but I'm not sure where you are trying to angle towards. There are a lot of ways to format code, and I can't tell you what works best for you.But this isn't shaking the paradigm of having an initial block of data (here separated a little awkwardly into two files) that is being subjected to a series of functional transformations, which is sorta pipe-like. Being able to talk about this as a 'pipe' is hugely useful, and having a pipe operator really standardises how people can talk about it. In my experience, selling this as a concept to a beginner is a lot easier when it has syntax support so they know they are doing the 'right thing' and get error messages/uncomfortable looking code when they try to cheat. I don't really care how you nest the brackets, if you want to write something that isn't a pipe go with (edit obviously a non-buggy version of...):
for i in 1:nrow(df){
if(df[i,'variable a'] > df[i,'x']) continue;
df[i,'variable ac'] = df[i,'variable_a']/df[i,'variable_b']
}
If someone shows me that and I can tell them to 'go rewrite it using pipe operators, it will be better' they will naturally start seeing that that horrible block of mixed ideas can be separated out into a few basic composable operations, and that is funnily enough easier for a beginner to grasp in my experience teaching people to use R. Plus, they can tell that a statement is going to be a series of composes on top of raw data if they see a pipe operator. This is a useful thing that makes it easier to read the code.Teaching a "paradigm" can be too limiting. Looking at some random tutorial on the web:
"To demonstrate the above advantages of the pipe operator, consider the following example.
round(cos(exp(sin(log10(sqrt(25))))), 2)
# -0.33
"The code above looks messy and it is cumbersome to step through all the different functions and also keep track of the brackets when writing the code.
"The method below uses magrittr‘s pipe (%>%) and makes the function calls easier to understand. sqrt(25) %>%
log10() %>%
sin() %>%
exp() %>%
cos() %>%
round(2)
# -0.33
Really? Do we want to teach people that the code below is so much better than the code above?What do we expect them to do if they find something like
100*exp(cumsum(0.6*diff(log(STOCKS))+0.4*diff(log(BONDS))))
which is a perfectly readable way of calculating the evolution of the value of a 60/40 portfolio of stocks and bonds from the value of each component at each rebalancing date?http://thatdatatho.com/2019/03/13/tutorial-about-magrittrs-p...
It’s easy to find examples of using the pipe in ways that I would consider suboptimal. But that doesn’t affect my thesis that, on average, the use of the pipe leads to more readable code.
I was also initially quite sceptical of the pipe, since it is a fundamentally new syntax for R (although obviously used in many other languages). I think the uptake by a wide variety of people across the R community does suggest there’s something there.
I had a similar experience with Lisp syntax: Clojure's threading macros seem a neat idea but I do actually prefer old-style nesting of function calls.
Maybe my R journey is a bit atypical. I started learning R around the time you created reshape and ggplot and used them extensively. But as the "tidyverse" thing has evolved I have found myself more attracted to "base" R as I've become more familiar with its data structures and functionalities.
df <- rbind(read.csv('/path/to/data/file1.csv'),
read.csv('/path/to/data/file2.csv'))
df <- df[, df$variable_a > x]
df$variable_c <- df$variable_a / df$variable_b
or maybe: df <- subset(rbind(read.csv('/path/to/data/file1.csv'),
read.csv('/path/to/data/file2.csv')),
variable_a > x)
df[,”variable_c”] <- df[,”variable_a”] / df[,”variable_b”]
Whether you think these are badly written pipes or not pipes at all, I don't care. It's good enough for me.In data.table, chaining operations together is standard and as simple as a set of brackets to contain the next operation.
Both are preferable to pandas though (joke not troll).
I have been using Python a lot longer than R, but the Pandas syntax is taking longer to internalize.
Maybe it's just me?
my_dt %>%
.[i, j, by]
It's a powerful combination. Fast, very readable, extremely versatile.One of the major innovations of the tibble package is that it puts the types of each column into the headings when it prints the table out. Dr. Matloff is grossly underestimating how painful it is for beginners to, eg, figure out when they are dealing with a column of strings vs a column of factors. If a beginner copies code to someone else for help, the other person can actually tell what types are involved! It is like magic compared to base R.
Given how ad-hoc the R community is and the paucity of people with programmer backgrounds, these minor usability improvements on controlling types are important. And anybody who cares about efficiency is using the wrong language.
EDIT Also, this comment that RStudio is bypassing an Open Source project's leadership - that is one fo the major features of open source. It is usually a healthy sign that the project leadership are lost in the weeds. I recall a situation with broad parallels in the history of GCC for example; it didn't kill the project although I believe a bad codebase died.
This is just sapply(df, class)... Is this not found in R 101 material that beginners are likely to find?
So to answer your question, no, its not just simply that.
- (l)apply: List apply always returns a list
- (s)apply: simplify apply tries to return simplified result
- (v)apply: verify apply checks the return type conforms to user supplied example
- (m)apply: multiple apply applies FUN to multiple vectors
- (r)apply: recursive apply is essentially a flatmap
- apply : no device here, only use on matrices, never data.frames
As much as I think people can lean too much on the concept, single responsibility principle would've gone a long way here.
That isn't going to change, because I'm just not going to use them or 'invest' the time in finding out what some statistician-of-yore's interpretation of a map is. Instead I'll stick to tidyverse map - returns a list. Or tidyverse map_[int/chr/dbl/etc, etc] if I want a vector of [int/chr/dbl/etc, etc].
That and the data frame manipulation verbs covers the most useful 80% of cases where *apply would otherwise be needed. If the base R team were implementing functions that way in the base, stats professors wouldn't need to complain about mass exoduses from base R.
I develop R packages for my colleagues and so I stick to Base R whenever possible. I don't want my packages depending on the tidyverse at all. But for EDA, I am agnostic about what my colleagues do. They should use the tools that stay out of the way and let them get their hands around the dataset intuitively. For me that's Base R, for others it's data.table or tidyverse.
There is nothing complicated about what sapply does... It simply means loop over the elements of the first argument (which must be a list; btw a dataframe, df, is internally the same as a list) and apply some function. lapply does this and returns a list, sapply does this and can optionally "simplify" the results into a vector, etc.
So:
lapply(df, class) = loop over the elements of df and tell me the class, return this in the form of a list
sapply(df, class) = loop over the elements of df and tell me the class, return this as a character vector
This is basically lapply:
res = NULL
for(i in 1:length(df)){
res = append(res, class(df[i]))
}
return(res)Its good to know tools like data.table exist though, people shouldnt think the tidyverse is the only way to do things.
Pretty much everyone knows the Tidyverse is slow to run, but many of us find it faster to write and very readable because it avoids one-time-use assignments (i.e. in `a <- f(x); g(a)` the variable `a` is only used once). Naming things is hard. Prof. Matloff's point is well taken though - you don't need the Tidyverse and it doesn't run faster. Fair criticism.
> The "star" of the Tidyverse, dplyr, consists of 263 functions. While a user initially need not use more than a small fraction of them, the high complexity is clear. Every time a user needs some variant of an operation, she must sift through those hundreds of functions for one suited to her current need.
> By contrast, if she knows base-R (not difficult), she can handle any situation.
Yet base R has some 1,346 functions. Another 455 functions in `stats`. I'm not sure the quantity of functions has as much bearing on learning as the quality and consistency. So this seems like a silly argument to me.
Mostly I think Prof. Matloff is forgetting that base R is kinda terrible and no one else has done much about it aside from RStudio.
I think he'd disagree with the contention that base R is kinda terrible. Like, I'm pretty sure he's wrong. But this isn't something he's forgotten, it's something he's never believed.
The point about the Tidyverse and RStudio is that RStudio doesn't make any of its revenue from the Tidyverse. They distribute it free. There are no upsells for any of it -- I can't go out and buy ggplot2 add ons, I can't buy dplyr support contracts. There's no money to be made on the Tidyverse, and plenty of resources thrown at it.
So what is RStudio selling? They sell an IDE, they sell servers to host things built in R, they sell a package manager solution, they offer most of these things as a SaaS offering as well (although rstudio.cloud is still in trial and they haven't started charging yet).
In other words, RStudio offers tooling for people who write R. They don't really care what kind of R you're writing, from a financial perspective. They get paid the same if you're using RStudio to write stuff using data.table as they get paid if you use dplyr. RStudio is making a bet that more people will want to write R code if they make R a more pleasant language to work in. And anecdotally, they're absolutely right -- I tried to learn R before the Tidyverse, and I hated every moment of it. I like post-Tidyverse R (although I got into it when people still called it the Hadleyverse). I know several other people with similar experiences to mine.
I enjoy base-R solutions where I can. I use data.table when my data sets are medium-ism. But like most people, my data is small and people just call it big and thus the attraction to dplyr in particular. Those chained pipes are addicting.
Always love seeing R discussion on HN, seems to be taboo though (not a 'real' language and all that).
- is based on scheme, a HN favourite;
- integrates very well with C++ through RCpp allowing you to do whatever you want. 9/10 times someone already went through the trouble for you.
I think it’s just insecurity. They worry a lot about whether they’re a good enough developer, so they try hard to be on the “most advanced” platform they can figure out, and then need to talk down to the other platforms.
Secure, talented, older developers, in my experience, tend to be open to whatever technology is most appropriate. They have confidence in their ability to learn new things. And see the good traits of any tech, as well as the bad one—a necessary skill if you’re going to be the kind of developer who can make traction in any situation.
The platform fetishists can only really work from inside their special tower.
For instance, why does R use <- instead of = for assignment? Because initial versions (of S) predate C -- developed down the hall -- the language that first introduced = for assignment.
Also, it's unbelievably slow for basic operations like looping, function calls, and variable assignment. It's literally orders of magnitude slower than Python (I've tested it). Unlike its fellow C-flavored-Lisp Javascript, R retains an extreme level of homoiconicity, Which apparently makes the language very difficult to compile or optimize. I don't know how e.g. SBCL does it, but over the lifetime of the R language no one has managed to implement a non-trivial subset of R that performs better than GNU R (fastR was never finished to my knowledge). So you are basically relegated to writing code in C, Fortran, or C++ if you need even decent performance. Otherwise you are stuck using "vectorized" operations like lapply(), which are fine but make for a jarring experience if you're coming from other languages.
So of course it's a "real" language in the sense that Bash is a real language. But it's ultimately and fundamentally a domain-specific scripting language and I don't know if there is a way around that.
Edit: looks like it's still a full major version behind? To be fair, there weren't many significant changes between R 2.x and 3.x.
I have not tried it in many years and I don't know how much faster or compatible it is. Nevertheless, I think it's a good thing that an "outsider" is looking at performance issues.
As far as I know it's essentially a one-man's effort. One could imagine that RStudio or Microsoft would be able to improve R performance substantially if they wanted to.
But for its domain, I think R is great.
I guess when all you have is a hammer...
It seems the Graal people managed to make R quite fast and is still being developed. Is there another faster? Any way, check out my blog post:
http://bootvis.nl/fastr-sudokus/
However, fastR is not a drop in replacement yet.
In any case, for me professionally, R slowness is never a problem. data.table does the heavy lifting.
What is a confusing mess, is the object system.
Then you can optimize and compile. This is particularly true of Scheme which foregoes much of the dynamic quality of Common Lisp in favor of a much more static view of the world. But its still basically true in CL.
R doesn't have macros of that kind. It has a sort of weird laziness based on arguments remaining unevaluated until needed. Metaprogramming in R is typically done by intercepting those "quoted" arguments and then evaluating them in modified contexts (this is exposed to the user more or less by letting them insert scopes into the "stack").
Thus, there is no distinction between macroexpansion and execution time. Hence, its tough to write a compiler, which is basically just a program which identifies invariants ahead of time and pre-computes them (eg, lexical scope is so good for compilers because you can pre-compute variable access). Because all sorts of shenanigans can be got up to by walking quoted expressions and evaluating them in modified contexts, R is hard to compile.
This is, by the way, why the `with` keyword was removed from Javascript. It provided exactly the ability to insert a scope into the stack used to look up variable bindings.
The idea of even the standard Common Lisp is that both is possible: a static Common Lisp and a dynamic Common Lisp, even within the same application in different sections of the program. Common Lisp allows hints to the compiler to remove various features (like fully generic code being reduced to type specific code), it allows a compiler to do various optimizations (inlining, getting rid of late-binding, ...) - while at the same time other parts of the code might be interpreted on the s-expression level.
The last thing I want to be thinking about is whether my compiler needs a "hint" that something can be stack allocated, for instance. That is their business, not mine.
A domain-specific rather than general-purpose programming language? Most definitely. A scripting language? I would dispute that. Scripting languages are for automating the execution of sequences of tasks that you'd otherwise run independently.
Function arguments/environments are still pairlists: a type of cons cell.
You can manipulate the parameters to a function as symbols before you evaluate them. This is useful — many very good R libraries all make heavy use of it — and certainly very interesting. Lispers probably know it as quote/antiquote.
I suppose we could also point that, much like most Lisps, R gives you plenty of object systems to choose from...
And that's about it for R-as-a-Lisp.
It's also a really bad language for working with tree-structured data, something Lisps normally excel at.
I’m sure someone will chime in that there are different runtimes, like Microsoft’s (does this run on MacOS? If not, how can I use it to develop and test on my laptop) or some expensive RStudio solution.
The fact is you have none of those problems in e.g., Python. Although I would prefer to use Caret for many predictive modeling tasks, it is non-trivial to take the R runtime to production. That was what I walked away with.
I doubt that anyone is going to be denied a job because they are an amazing R programmer but they just don’t have the experience with a particular set of packages, especially those that are as well implemented and easy to learn as the Tidyverse.
I can totally imagine someone not getting a job because they are a crap programmer, and then blaming it on something else. Or I can also imagine someone saying that the Tidyverse packages suck, and then be denied a job because of their attitude.
I could make a similar argument for most of the substantive claims in his post.
Some of these examples (dplyr vs. data.table) are cherry picked. I have several of my own examples where read_csv is way faster than read.csv, so maybe, like all good programmers, we should be testing and profiling our code and implementing the parts that make the most sense for our needs.
The bottom line is that the Tidyverse is a good set of packages and you are free to use them or not. There isn’t some blood feud (like vi vs. emacs) between users and non-users, we all get along just fine.
There are dozens of great tutorials on how to learn R that don’t use the Tidyverse, and RStudio is under no obligation to offer a full course on every possible way to learn R. Any R user is also a competent Google user.
It’s fine to be opinionated. But, a professor as respected as Norm Matloff should be careful of how they say things, or they risk souring their students on a set of packages that might be very useful in the future.
The problem is that dplyr is much slower than data.table and RStudio is promoting tidyverse too much to make the slower choice a default for many users.
The article is 100% correct on this issue.
dplyr syntax is definitely more concise and readable than base R, but comparing to data.table I don't think it has any advantage in terms of saving time writing or reading code.
https://stackoverflow.com/questions/21435339/data-table-vs-d...
I feel like the first example is far more readable than the second. People can disagree on this, but the adoption rates of dplyr versus data.table do suggest (don't prove, but suggest) that the consensus on the issue leans towards dplyr. As we've noted, people certainly aren't adopting dplyr for the speed.
diamondsDT[cut != "Fair", .(AvgPrice = mean(price),
MedianPrice = as.numeric(median(price)),
Count = .N), cut][order(-Count)]
There is no need to break it to 10 lines.And I don’t expect RStudio to say, “we have this collection, which works fine for most users, except in this case you should replace package x with y, or in this corner case you might like package z.”
RStudio doesn’t need to promote a fragmented ecosystem if they don’t want to, it won’t cause the death of R.
RStudio isn't on the leadership team. If the language ends up dying (which it won't, R is dug into it's place like a tick), it would be due to the leadership team's refusal to adapt to good ideas coming in from the community.
This article only briefly touches on what I think is the biggest issue with the Tidyverse. The Tidyverse is incredibly limiting. The "tidy" workflow hides many details of the language, which leads many users to think all data has to be in a tidy "data.frame" (or now tibble) and organized just so. The functions the Tidyverse authors have chosen to implement are seen as the limits of R's capabilities. Every data problem has to be forced into the Tidyverse box (or it is impossible). I cannot tell you how many Tidyverse scripts I have seen where 90% of the operations are munging the data into the correct format for the Tidyverse functions, when the original data can be handled with a single base R function (e.g. lapply). Most Tidyverse users accept this as just the way R is.
Most R users who only learn the Tidyverse never hit its limits, and for them the Tidyverse is perfectly OK. Those that do either resign themselves to the perceived limits of R, or (hopefully) start learning some base R. One friend who took the latter path once exclaimed to me "Logical subsetting is amazing!?!" - this is a foundation piece of the language. That she went 2+ years in the "Tidyverse" without even knowing it was an option was eye opening to me.
To its credit, the Tidyverse is empowering - with very little programming knowledge a beginner can do a lot of data science. However, the majority of these people get stuck as "expert beginners". Some of these become fierce advocates of the Tidyverse without truly understanding base R. Meanwhile, none of the truly advanced R users I encounter use the Tidyverse.
ed.typo
This is most likely because truly advanced R users have been using R since far before the Tidyverse existed. A whole new generation of R users is being brought up with the Tidyverse, so im curious to see how the situation will be in 10 years time.
That said, I have heard my mentor say "I don't get the point of <insert tidyverse function> - I did this 20 years ago."
Tidyverse has a tendency to promote "new" things without any references to what came before. If you listen to their talks they speak as if they invented functional programming and the idea of pure and tiny functions working together.
And it trickles down to the users. A lot of people who learned tidyverse first for example, praise the `purrr` package, but have no idea that something like `Map()` is in base R.
When I wrote code to do this, the preferred tidyverse routes would either be to 1:nrow(df) %>% map(function(index) { row = df[index, ] }) or else df %>% pmap(function(.... laundry list of variables here)) { }. A brief Google shows this has gotten worse in the last few years as dplyr deprecated rowwise operations. I do see stack overflow posts of people writing their own tidy/pipe-friendly row-wise iteration functions, but nothing official.
Maybe I missed something. I am a reasonably competent R programmer and package author, but I don't live and breathe tidyverse data-wrangling the way some people do.
If I understand correctly, you want to know how many NA's there are in each column in a wide-form dataset (as opposed to a tidy dataset)
# One line to make the data tidy.
# The form of data will be 3 columns: id, question, answer, and no, we don't care what the columns are called, except for id.
tidydf <- df %>% gather("question", "answer", -id)
# one line to do your check
tidydf %>% group_by(id) %>% summarise(n_NA = sum(is.na(answer)))
Tidyverse is highly opinionated about its data structure, and it is one of its limiting factors, as it basically treats every dataset as a sparse dataset. This actually fits very well with your data, as a datapoint is not a fixed questionnaire, but rather a datapoint is a respondents answer to a question (as questionnaires vary in questions, a tall table layout is quite fitting).From there on you have to think in groups and summaries, unless you wanna fight the library.
Tidyverse is an 80% datascience solution. It solves what you need 80% of the time really, really well, and the last 20% you either have to fall back to base R or really torture dplyr.
The tidyverse just 'made sense' to me when I started using R for the first time a few years ago, and now I love using R and programming. On the other hand, some of my ex-classmates learnt base R (because that's what we were taught) and found it hard, didn't learn anything properly, and now still think R or other programming languages are opaque and hard.
I'm not particularly fussed if Statistics Profs prefer data.table to dplyr or base R to tidyr, I know what is easier to teach, understand and use for me and a lot of other ecology/bio students and people.
Programming in base R is more akin to assembly language and has accreted a babel of inconsistencies that make it difficult to teach and learn. Learning base R isolates you into a Galapagos island of academics who are either ignorant of the needs of data workers or too elitist to engage with those not in their priesthood.
Learning Tidyverse is a considerably better transition for learning other languages, frameworks, and libraries.
Functional programming is closer to algebra than indexing into data structures with magic numbers. I've found more success teaching functional pipelines of data structures using the idioms in Tidyverse as a general framework for data work than base R. Abstraction has a cost but for learning it is the appropriate cost.
I sense that much of this `monopolistic` fear mongering is really about feeling out of date.
I think this really depends on the end point. If you want to learn to read data into R and do basic manipulations, plotting, modeling, etc., the Tidyverse absolutely has a lower bar to entry. Once you get into writing functions, it gets a little trickier. Knowing some of the base R programming concepts and skills will make you much more efficient. If you start debugging and profiling code, only knowing the Tidyverse becomes a liability because you fundamentally do not understand R's computational model (the Tidyverse does not follow it). Hence, if your end goal is to write and debug functions in R, the steeper learning curve of base R can more than pay off. If not, then the Tidyverse's low bar to entry can be more attractive.
Imho, the data problem Tidyverse is trying to solve is basically the ones we face in a database. So, select, join, inner join and so forth. Show me all the rows in this datatable where the 4th columm is larger than the 6th column and the number itself is odd. Something like that.
There might be other ways to do it, but you want your select, filter, summarize, mutate etc functions to all work with each other, pipe to each other and be compatible.
Maybe there is a better way to do all this -- I haven't seen it but I am not an expert -- but you have to show that to me.
So, in base R, walk through a set of example of mutating, joining, filtering and so forth, and show me how they are all easier. Then I'll say, wow there is an alternative to this Tidyverse thing. But in lieu of that demo, this felt more like an intro to a complaint than an actual complaint.
Edit: Also, its funny that Wickham is (apparently) such a nice fellow that people go out of the way to be nice to him in critiques.
I run a data science group a large geospatial company and we develop day in and day out in R and python. We've purged tidyverse as much as possible from all of our code base. We've moved completely over to data.tables, which make the vast majority of the tidyverse irrelevant.
However, the claims the author makes are entirely valid. Data table is a fantastic package, and easy to learn. Tidyverse really has split the R community in disparate sets of R users.
To my mind, one of R’s greatest strengths is its meta programming capabilities. It is this that allows such broad paradigms to exist within the same language. In that, the challenges that R is presented with is similar to languages like Lisp dialects and Scala.
I don’t have solutions, but the author has convinced me that a problem is indeed there.
I tend to agree. I've found tidyverse code has a write-only quality to it. Since I'm going to see more of it I plan to dive into the inner workings of at least dyplr and purr.
That said it is hard to deny what Hadley Wickham has done for the R community and I can't write off the idea that he sees a bigger picture I'm missing.
Correct usage requires learning the right abstractions, but fortunately these abstractions are shared across language communities and frameworks.
If you are doing any data processing work and you do not know about foundational functional concepts ie map/reduce then I would argue the code you write is less readable, overly focused on the idiosyncracies of how to do something rather than how entities and data are related to each other - their logic and structure.
The main optimization is to be correct and expressive of one's intention and purpose. If you need performance use Julia.
So they end up with a subset of functionality they can get stuff done with, but then they need to collaborate with someone who uses a different subset. That can be messy.
I think in an environment where everyone learns R from the same course/book/MOOC/whatever (or have done lots of different sorts of programming) and the organization can impose a style guide the tidyverse approach would be great, but when you have people coming from all sorts of places and backgrounds I don't think it's a good fit.
What’s really crazy is I never realized how fast and easy excel pivot tables were to work with.
I know I know it sounds ridiculous.
But if you do a lot data splicing and dicing, excel can actually get your cuts out way faster through pivot tables than writing R code.
So if you use R for almost everything, give Excel a try as well. And the nice thing is this is even more powerful if you work with non data science folks, because then you can say I just did this in excel (where excel skills should be required by everyone in the company, and R for the data scientists) and they’ll probably just leave you alone the next time they need something since they will try to do it in excel first.
I tried my hand at saying something more original about this debate recently: https://unconj.ca/blog/for-and-against-data-table.html
On the other hand, it's nice to see some criticism of RStudio and Hadley, even if it does stray into the conspiratorial. They've done a lot for the R community, but they do have some obvious blind spots.
The original article is more about dplyr though.
Data.table and data.table-esque notation represent such an improvement of tibble/dplyr, that within my company, we're making a concerted effort to purge all tidyverse packages from general use (less ggplot). When new developers come on, if they are coming from tidyverse, their first task will be something involving pipes and data.table. Tidyverse was fine in school. It doesn't pass muster in production, at least not in our work.
Data.table syntax is simpler, easier to read, easier to teach, and orders of magnitude faster. It plays nicer with other packages than the tidyverse (if it fits into a DF, it almost always fits into a DT, and i've never met a tibble that I didn't wish was a data.table), and since almost all of our datasets are 10's to 1000's of millions of lines long, the decision was really made for us.
This is a bit disrespectful.
"Data.table syntax is simpler, easier to read, easier to teach..."
This is rather arbitrary, and I don't think it's the majority view of the community, whatever the advantages of data.table.
"...orders of magnitude faster"
This is an exaggeration in most real cases, even according to the benchmarks pointed to by data.table.[1]
My greatest regret about coining the word tidyverse is that for some reason people seem to think it’s a monolith. It’s not; you’re totally free to pick and choose whatever parts of it you find useful.
It doesn’t hurt my feelings if packages that I have help write aren’t the perfect fit for your problems. Use whatever makes you happy :)
I probably should have phrased my comment better. My point was that it’s perfectly possible to use ggplot without explicit knowledge of tidy principles or tibbles. This is great! And for example, I’m reading through your ggplot book and it doesn’t make reference to it.[1] In my work, we use data.tables (we believe we need the raw performance) with ggplot for visualization and are not using the rest of tidyverse and it works fine.
[1]: In recent versions this seems to be slowly changing.
But there’s no reason to use only the tidyverse. That’s not something I’ve ever recommended and it would be extremely hard. I just object to people claiming that some of the most important parts of the tidyverse aren’t actually parts of it.
If you want efficient code, try not to use dplyr. If you want readable code, data.table isn't the best answer. If you want speed, there are better languages than R to choose from, even with Rcpp.
If you want to get something done, go with whatever you know best and will do the job.
Avoid Stata.
Check it out: https://www.tidyverse.org/articles/2019/06/rlang-0-4-0/
However, I'd claim most folks don't deal with big data or even medium data! [Aside: I chuckled at the 1e5 columns on the plot]. Personally, at the small-to-medium range, all languages are fast enough. It boils down to ergonomics for me. I _detest_ the subscript notation. It slows me down -- not just productivity-wise but also how I read or reason-about my code. Not surprisingly, I prefer `dplyr` over `data.table` (base R or `pandas`)
compare base R:
var1 <- "pounds"
var2 <- "wt"
mtcars[[var1]] <- mtcars[[var2]]/1000
Now hadley solution comment: You can now express that sort of thing fairly elegantly with dplyr: var1 <- sym("pounds")
var2 <- sym("wt")
mtcars <- mutate(mtcars,!!var1 := !!var/100)
Watch your steps: !!, :=, sym.
Thats a lot of added complexity to be able to use column names, and to top it all, the use of non standard evaluation make things very complex when one go a little away from basic EDA.mtcars %<>% mutate(pounds = wt / 1000)
Which is obviously simpler, and quite likely what a beginner is actually trying to do.
The real complaint - "tidyverse doesn't let us use variables to name columns!" - is fair enough, following quasiquotation is hard work. If you need that, drop back to base R. However, it is harder to teach because all the common operations (select, filter, etc) will involve some combination of []/[[]] and relatively hard to figure out which one a piece of code is trying to do.
Except for the lucky few who have brains wired to think in terms of arrays and database relations, someone learning the language is unlikely to thank you for that.
I have been using data.table from almost day one, it does worth more recognition. The syntax is a little bit hard in the beginning but can be grasped after some efforts. I often feel that my exploration code will be much more tedious look if written in dplyr style.
It's, like, the easiest data manipulation library in any language.
I'd like to understand better why the data structures work the way they do and thus have an intuition on what operations to use when. The O'Reilly Python Data Science Handbook[0] seems like it might be useful here, but I'm not sure if it is still up to date.
[0] https://www.oreilly.com/library/view/python-data-science/978...
The OReily book should not be that outdated if at all. It also covers other essential tools.
[0] https://pandas.pydata.org/pandas-docs/stable/getting_started...
In fact python is a terrible language and the only reason anyone should use it is for access to sklearn, scikit, pytorth, and pandas.
Hopefully Julia will be able to unseat python for data analysis in the future.
I was (sadly) amused by a recent comment on HN linking to some tweets celebrating that some R code still ran four years later.
Maybe I lack context, but those tweets looked to me as if someone said “I can run this four year old game in the latest release of Windows!” or “Three versions of Office later I can still open my excel files!”