The case for a pipe assignment operator in R
hughjonesd.github.io
hughjonesd.github.io
Even if the piped argument goes in a certain location by default, it still creates an exception to all the other arguments' syntax, which adds another layer of complexity.
I guess at heart lisp syntax has always made the most intuitive sense to me personally. algol/c does as well, but the more you mix them the more I tend to dislike it.
I guess pipes always feel like you're "proceduralizing" functional constructs, which makes things more murky (although maybe that reflects some deep equivalence between the two?)
This is just my personal preference, and in an R context. I could be convinced differently. Part of it too is I feel like R is fragmenting by lots of defacto standards, and not in desirable ways. I increasingly find myself wanting to use other languages, even more so than in the past.
They do this by introducing new idioms. Granted they are idioms that if every r user would adopt 100% and all older code were rewritten 100% would in fact make the language more tidy. But that will never happen. So it ultimately is counterproductive imo. If I want to write concise shell style pipelines for data analysis I could use awk. If I wanted readability I could use python. Don’t try to cram all of those things into Posit/r.
a <- a |> mean()
Allows me to test `a |> mean()` to check it works before it gets assigned to replace a. Very useful when you have very long pipe commands. <|> Wouldn't allow this.
x %>% set_names(., str_c("X_", names(.))
All the underscore does is let you pass the piped argument as a named argument instead of the first otherwise-unfilled argument.y = h(f(g(x)))
where do you put the print statements or loggers?
y = h (f (traceShowId (g x)))
(Haskell is normally pure, but these functions from Debug.Trace are a notable exception.)Or use Stickers provided by Sly which sort of log things as they change. https://joaotavora.github.io/sly/#Stickers In that case you'd highlight the form starting at f, and turn a sticker on. Then run the code, sticker outputs into a window.
This release also cleaned up some edge cases so that the semantics are as similar as possible too, making it easy for folks to switch between the two.
And then it all made sense. Non-coders equate R with RStudio to the point that they don't really get that R and its editor were separate concepts, and therein lies the power.
Everybody uses the same IDE and it always works. No more figuring out why the language server isn't offering completions or wondering why ESS has such awful defaults.
R (and the tight integration with RStudio) makes it very easy for non-programmer types to conduct their own sophisticated exploratory analysis and to produce good looking plots without mucking about in Excel.
I still hate the `<-` assignment operator though.
1. Unmatched simplicity of dplyr for data manipulation.
2. Great graphs with ggplot2.
Oh, and 1-based indexing. Try telling economics students why 3..5 means elements 4 and 5.
Doesn't 3..5 mean elements 3 and 4 : the 3rd and the 4th? 0-based indexing might feel less intuitive actually (element 3 is the 4th, and element 4 is the 5th). I don't know much R, I'm assuming that in the n..m notation, n is included and m is excluded (edit: also included, see replies).
I don't now about economy, but I think mathematicians use 0-indexed and 1-indexed sequences interchangeably and they usually (have to) specify which is the first term (u0 or u1).
Yes, you’re agreeing! The point of the parent was that 0-indexing is less easy to grasp
var
arr: array [1 .. 5] of integer;
begin
for i := 1 to 5 do
arr[i] := i
end;
Normally you would write "for i := low(arr) to high(arr)". I just wrote the numbers "1 to 5" to be explicit here.That being said many decent GUIs have been developed for a more traditional analysis feel, but still use R under the hood, I like JASP the most out of the current options.
"Assign X with value equal to 5"
Personally, it seems to me that = for assignment is so ubiquitous that this argument has become stale.
I still use the arrow because it's the prevailing recommendation in R style guides.
Admittedly this kind of goes out the window with a lot of tidyverse non-standard eval stuff. E.g. if you do `df %>% mutate(x = 5)`, then x does get set to 5 in the data frame, so that is logically an assignment of x in a visible scope using '='. Going by my logic, functions like mutate should ideally use a different symbol for operations that alter the input like this, but we run up against syntactic limitations of the language.
Why on Earth did you think that? Because it's not what the hip frontend devs use?
I get it, you like R. I'm not interested in litigating this any further.
mtcars |> filter(am == 1) |> select(mpg, gear) -> df
I don’t see every using a read and replace pipe operator over the right assignment approach.
The evidence section kinda lost me (what is `wardmap@data` supposed to be? - I'm not sure it's valid R code, and can't spot it on stack overflow as the article suggests?)
Small correction: under the bit about 'A simple syntax transformation converts .. into ..' the two should be the other way around.
An assignment pipe should be a relatively easy sell since it produces more concise code, and lets the developer work left to right, top to bottom (how we all want to work), rather than having to backtrack to place `df <- ` at the start.
I'd try demo it on super common use cases (so the average R user can easily see how coding patterns are improved by it).
Found the Stack Overflow question: https://stackoverflow.com/q/34500567
EDIT: came across a nice explanation of the S4 system [1] under the 'Overview' and 'Slots and accessor functions' sections. Seems it's commonly used in genetics and geospatial work.
[1] https://kasperdanielhansen.github.io/genbioconductor/html/R_...).
I did it just to avoid the dependency because truth be told the native pipe seems a bit ugly to me (but that’s just me)
pandas tries to chain methods but it isn't the same and .apply() is ugly.
This <|> I'm not sold on, though.
It feels like R is in decline though.
mydata <- read.csv('mydata.csv')
mydata['score'] <- mytransform(mydata['score'])
This can be simplified with your proposed operator: mydata['score'] <|> mytransform()
An upgrade in elegance, i like it. But in R, a language which is commonly used switching between scripts and the REPL, with an IDE which by default captures your workspace so all your variables etc are restored the next time you resume the session, in this environment i feel like variables should be used as constants as much as possible. Mutating a variable (as in my example) creates room for confusion: Does mydata contain the raw csv data or the transformed data? Did i already evaluate this line in the REPL or did i not? What happens if i evaluate this line twice? Many people i know tend to "jump around" in their scripts, not following the written order of operations. This creates potentially irreproducible environments.
My proposed solution is treating all variables (or at least as many as possible) like constants: mydata <- read.csv('mydata.csv')
mydata_transformed <- transform(mydata)
Now i always know what a variable contains because each variable contains exactly the same value, independent of when i evaluate.But nevertheless; i kinda went on a tangent here and strictly speaking, the problem i describe only arises from careless user behavior (which is quite prevalent in statistics though). Aside from this kind of behavior, i think this operator is an elegant idea!
R is used because it remains the single best language for data manipulation and reporting.
… |> mutate(foo = f(foo)) |> …
You can write df$foo <|> f
But that seems to not be the case? You need to introduce new names, e.g. … -> tmp
tmp$foo <|> f
tmp |> …
Which seems worse to me. Maybe I’m missing something.For comparison, consider jq's |= operator where you might write e.g.
… | (.foo |= f) | …
Which I think composes better with pipelines that do not “mutate” (or rather, make substitutions) …|> (\(x) x[vars] <|> bar())()
But now we are getting into self-parody… as the article says, it’s a shame there’s no nice anonymous function call syntax.Or possibly
… |> within(foo <|> bar()) |> …but actually i quite like R for graphing and such stuff, despite wrestling with the syntax.
I am hopeful of S7 but I still believe that R needs some proper PL engineering.
It has many dangerous corner cases and mutability when it shouldn't.
I migrated to Julia for my experiments/computations. Safer.
Still R has some great ideas but not developed to their proper extend.
S7 is in the right direction. Satic typing could also be a big thing.
The comparison in this example is misleading:
# before
names(data)[1:2] <- paste0(names(data)[1:2], "_suffix")
# after
names(data)[1:2] <|> paste0("_suffix")
Since we're talking about pipes, the first option should use pipes: # before
names(data)[1:2] <- names(data)[1:2] |> paste0("_suffix")
# after
names(data)[1:2] <|> paste0("_suffix")
For pipe enthusiasts this is already pretty clear. The thing on the left and right side of the assignment is the same. Compressing this saves characters and avoiding repetition of the 1:2 part is nice, but I don't know if the cost in familiarity is worth it. In any case involving a data frame, I would prefer using the .cols parameter of the rename family of verbs to either of these base R approaches.Also, a linked post by this author [1] is materially outdated on the use of case_when and across (of course I sympathize - the tidyverse has moved very fast in the past few years). Thanks to the new backslash notation for anonymous functions and various other dplyr upgrades, it is very elegant to do assignment across multiple columns based on conditions evaluated against the entire data frame. Behold:
mtcars %>%
head()
mpg cyl disp hp drat wt qsec vs am gear carb
1 21 6 160 110 3.9 2.62 16.5 0 1 4 4
2 21 6 160 110 3.9 2.88 17.0 0 1 4 4
3 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1
4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1
5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2
6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1
mtcars %>%
mutate(across(mpg:hp, \(x) case_when(wt < 3 ~ x * 10, wt >= 3 ~ x))) %>%
head()
mpg cyl disp hp drat wt qsec vs am gear carb
1 210 60 1600 1100 3.9 2.62 16.5 0 1 4 4
2 210 60 1600 1100 3.9 2.88 17.0 0 1 4 4
3 228 40 1080 930 3.85 2.32 18.6 1 1 4 1
4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1
5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2
6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1
Over the past decade, one thing I have learned is to never, ever count Hadley Wickham and the tidyverse team out when it comes to optimizing their API. If there is a lack of expressiveness or orthogonality, they will fix it. The R community will take a while to absorb their ideas because there are so many idioms flying around (partly the fault of previous versions of the tidyverse), but the model they've landed on recently is incredible and should be a model for tabular data manipulation in any language. data <- data |>
filter(…) |>
mutate(…) |>
left_join(…)
… Etc. This would simplify to data <|> filter(…) |> …> data <- data |> > split_by(field) &> > # mutate over all > mutate(calc=...) 1-5> > # filter the first 5 elems > filter(…) &]> > # control back to parent > collect() ]> > # collect all elements and > # implicit rbind them > ...