R 4.0
stat.ethz.ch
stat.ethz.ch
df_iris <- iris
tb_iris <- tibble(iris)
nunique <- function(x, colname) length(unique(x[,colname]))
nunique(df_iris, "Species")
> 3
nunique(tb_iris, "Species")
> 1
Imagine now using some complex function from a repository (i.e. BioConductor) that works on data.frames and passing a tibble to it.Original post follows:
Right off the bat, the problem is not "using" tibbles, it's that you've incorrectly constructed one by passing the data through the tibble() constructor rather than using as_tibble(). The tibble constructor -- for pretty good reasons in other circumstances that seem crazy to you here because of your intent -- infers that you want the entire data frame to be a single column inside the tibble, called "iris". It does this because it evaluates the variable name passed to the tibble constructor as both the intended column name and the data to be placed inside the column. This demonstrates nesting, which is one of the great features of tibbles and otherwise used for a bunch of stuff.
If you had done `tb_iris <- as_tibble(iris)`, it would have worked fine. `as_tibble()` is the function to convert an existing data structure to a tibble. R is obviously not "type safe" in any way, but you can engage in defensive programming, and one way you can do that is being hyper-aware of the steps you take during type conversions. If you check the documentation for `tibble()`, it tells you explicitly to "Use as_tibble() to turn an existing object into a tibble." Is there a reason you didn't? Imagine this related example:
my_string <- "10"
numeric(my_string)
as.numeric(my_string)
Would we conclude that "using the numeric type can be dangerous" because the constructor interpreted the argument different than the conversion helper?Second, I suspect you must be using extremely old versions of things, because on more recent versions, your nunique function would fail, not produce 1. I correctly get "Error: Can't find column `Species` in `.data`." This error message is maybe a little confusing if you don't check the structure `str(tb_iris)` of tb_iris to see what I mentioned above, but is the correct error to output in light of it. You'd also be able to flag this by just checking `colnames(tb_iris)` or `View(tb_iris)` if you're working in RStudio or using the embedded environment pane or really any other way of looking at the data.
But your broader point is also false. Once a tibble has been formed, it should work EXACTLY the way a data.frame works because R objects can have multiple classes. The only thing that makes a tibble different than a data.frame is that it has an additional class label. All dispatches that work on data.frame objects work on tibbles because of how multiple classing works in R. This has been a goal since the beginning of tibble. The one exception I'm aware of is external functions that incorrectly check `if(class(obj) == "data.frame")` instead of using `is.data.frame()` or `if("data.frame" %in% class(obj))`. The former is and always has been incorrect because of how multiple dispatch is designed to work in R and should generate an error with multi-classed objects because the if statement evaluates to a vector of logicals instead of a logical.
Once way you can tell that tibbles and data frames are identical save the above caveat is to run the following code:
df_iris <- iris
tb_iris <- as_tibble(iris)
identical(df_iris, tb_iris)
class(tb_iris) <- "data.frame"
identical(df_iris, tb_iris)
Note that you are not "downconverting" a tibble into a data.frame in this code (but that would work too) -- you are taking the tibble exactly as is and hacking its class label to look like a data frame. It's identical because a tibble was always a data frame.First, about the as_tibble - it returns the same thing as tibble:
tb_iris <- as_tibble(iris)
length(unique(tb_iris[,"Species"]))
> 1
Second, about the incorrect version: > packageVersion("tibble")
[1] ‘3.0.1’
Which is also the current version on CRAN.Third, about the classes:
You say:
> Once a tibble has been formed, it should work EXACTLY the way a data.frame works because R objects can have multiple classes.
This is not the case. You can add any class to any object in R S3 system. So people behind tibble can call their tibble a data.frame but it gives no guarantee that it will behave like one.
More about this problem here (and you can also find replies from tidyverse authors) https://stat.ethz.ch/pipermail/r-package-devel/2017q3/001896...
I highlighted the nesting issue in constructing versus coercing (which is correct and does have implications for what you're trying to do) but actually in your example the distinction is broken because of a different edge case
Which is to say the following:
ncol(iris) # 5
ncol(as_tibble(iris)) # 5
ncol(tibble(iris) # 1
iris$Species # Works
as_tibble(iris)$Species # Works
tibble(iris)$Species # Errors because of nesting
iris[, "Species"] # Works
tibble(iris)[, "Species"] # Doesn't work
as_tibble(iris)[, "Species"] # Works
However, you're correct that because the subset operator for tibble doesn't drop dimensions, length gets you the number of columns rather than the number of observations. This does speak to the fact that length is a pretty shitty function to begin with, but I concede you're partially correct there.You are also correct that because class labels are not contractual, there is no guarantee that having the data.frame fallback label means stuff behaves identically (for instance, you could add the data.frame label to any data structure and the data.frame dispatch stuff would not work properly). My point was that in the case of a tibble, a tibble is literally a data frame with an additional class label. If you remove that class label, it's exactly identical.
But your example and linked discussion does highlight a way in which I'm wrong; the subset function is overridden for something with a tibble class label. That's true and could produce edge cases I hadn't considered.
Apologies for any hostility in my original reply.
The other OO systems in R do act closer to traditional classes, but all the tidyverse stuff is S3.
(But the OP was correct in another sense related to the example narrowly!)
I personally think it's a good thing that the drop-argument defaults to FALSE for tibbles, since data frame's default drop = TRUE is a source of frequent bugs. The change of the default for this parameter is the source of your observation.
One might ask whether it was a good idea that tibble enlists data.frame as an inherited class. Since a tibble obviously doesn't behave like a data frame, one could also argue that this is a mistake on part of the tibble developers but this is a different discussion.
As for whether or not tibbles should be data.frames - I posted a link to this exact discussion on R-dev mailing list within this thread, as an answer to a different poster. Here it is: https://stat.ethz.ch/pipermail/r-package-devel/2017q3/001896...
The drawback is that character takes up so much space, but these days memory is so bountiful it usually doesn't matter.
options(stringsAsFactors=FALSE)You can always run R --no-init-file to be sure that you have the default settings. Now you have to know what default settings the code that you want to run expects.
If you're happy with Python, by all means keep using it. I use both languages. Just suggesting that if you gave up on R that long ago, you might be pleasantly surprised by how much better it's gotten since then.
That said, R still is a stats language by design. In Python or JS for example, you can concatenate strings with ‘a’ + ‘b’ but the + operator in R is explicitly only for numeric types. R also has a horrible architecture for memory management, leading to code that uses profound amounts of RAM. I face this issue constantly as I work with very large datasets.
I have a love/hate relationship with R and despise using it in production. I’m also not a fan of the divergence that Tidyverse has caused. Particularly the expectation of Non-standard evaluation and the tendency for new R learners to become dependent on these packages. Especially as it relates to reproducibility and deploying code, these unnecessary dependencies suck. Tidyverse is maturing and breaking changes are still too common for comfort. There is no reason in my opinion to load stringr when a grep() will suffice. Or, to subset with select when [[ works perfectly fine. Or to filter when subsetting on a logical with which()... the list goes on. Tidyverse is essentially reinventing the wheel in many places. The biggest problem is that it doesn’t translate well to base-R in my experience with new programmers, leading to this divergence.
That said, piping with %>% and modifying directly with %<>%(via magrittr) is a pleasure that other languages I’ve worked with don’t manage as well.
And at the end of the day, I’m not going to rewrite implementations of all the latest statistical methods already written in R, and this is its strong suit. I’m increasingly using sophisticated spatial and spatiotemporal methods, and these methods are solely implemented in R.
I understand that R gets a lot of flack from software developers and I understand why. But, I also think it’s too often overlooked for its strong suits.
Plus, it's not such a bad feature when you know it's coming.
As far as the tidyverse goes, I get it now however it seems to discourage the creation of a nice, well organized set of functions to limit the amount you need to keep in your head at the same time, and a lot of R users are very smart people capable of understanding very disorganized code. Instead of functions you get copy/pasted incantations, in Base R it's at least broken down into steps which is a start.
I teach base R first for a few weeks, then I teach the TidyVerse as I introduce data science concepts, like text mining. I convey the TidyVerse as like an overlay on R, improving shortcomings and adding great functionality, syntactic sugar, etc.
This context switch is jarring to some students. It's great to see some of the TidyVerse's strengths--part of the tibble--be moved back into base R.
Now pardon me: I need to go. I start teaching dplyr in 51 minutes!
I do exactly the same, and I always find it hard to justify base-R in front of my students, after they learn about the Tidyverse ;)
The TidyVerse is a opinionated group of R packages that all function similarly. ggplot2 is in there, and I already use that one.
I do like R for analysis. (I choose to use it over excel or Stata for my data analysis/ classes), though I wish the help pages were better.
Its a hard language to grok, even the data type names are unclear (data frame, vs matrix, vs vectors..). I'll take a look at the tidyverse packages.
https://www.tidyverse.org/learn/
any additional learning resources you'd recommend?
Select, for columns Filter, for rows Mutate, for adding rows or changing values
On top of that there are a bunch of sugared versions of these, like rename, instead of select or transmute instead of mutate.
When you want to do something advanced you either combine those or go out to see if somebody else has already made that function for you. I don’t think that having functions for joins or removing NA’s is something special to tidyverse. Everybody needs to put such operations somewhere.
I find tidyverse to have a simple, expandable vocabulary that is easy to learn, read and grow.
But it is slow.
I never found any kind of logic behind data.tables API, and I find it hard to read. I haven’t found an introduction to it that makes anything click.
But it is fast and it has these absolutely amazing rolling joins
This is how I learned dt.
How does selecting the 3rd column,
dt[, 3]
jive with>`DT[i, j, by, ...]` which means: “Take DT, subset rows using i, then calculate j ...
I don’t understand how column selection and calculating an expression is the same thing.
So now I can’t figure out if
DT[, sum(V1)]
Will return the column indexed at the sum of V1 or if it will return the sum of the column called V1. How can I tell, from the syntax?So for your first example it grabs the third column. The calculation you’re performing is a column selection.
For your second example, you’re returning all the rows, selecting column V1, and summing it.
I really prefer to name the operations
dt[rows, columns, groups]If I want to use an advanced data manipulation library, I'd typically reach for data.table. If I want to use a verb based approach, why not SQL rather than dplyr.
I have tried dplyr and code I've written a few years ago still imports that package (hopefully it still works), but I just didn't find it was particularly useful or helpful compared to the alternatives.
I learned R in a way that everything went right to left. The variable on the left was manipulated by whatever (functions or other code) on the right.
The pipe operator reverses that flow, with everything now moving left to right, which I find difficult to follow and debug.
data %>%
filter(thing > 4) %>%
mutate(new = fun(old)) %>%
group_by(var) %>%
summarize(new_mean = mean(new))
Then you are reading the code top-to-bottom, just like any other R script but without all of the temporary variables.It's great for data analysis and terrible for programming.
The reason it sucks for programming is because you can't debug it or inspect intermediate variables, which is really really annoying.
It's great for one off transformations and plotting, but a really, really really bad idea for programming.
Then you add NSE, which makes it hard to functionalise procedural pipes (especially for people who learned tidyverse) and it's a recipe for unmaintainable and profoundly annoying legacy code.
That bring said,I love it for interactive analysis.
As someone who taught myself base R from scratch in 2013, I (used to!) agree with this. When the pipe operator was first introduced, I’d roll my eyes whenever I saw a script that used it and move along.
But I forced my brain to adapt, and now it’s probably my favorite feature of the R language. Data science is full of sequences of transformations, and in my opinion it’s more readable and bug-resistant to phrase these long chains as:
f(x) %>% g() %>% h()
rather than: h(g(f(x)))
or certainly: foo <- f(x)
foo <- g(foo)
foo <- h(foo)
I can comprehend and modify others’ (to include past versions of me) R code much more quickly with this paradigm. You can quickly debug a chain by commenting out functions sequentially (i.e. first test: “f(x) # %>% ...”). It also becomes much faster to plug new transformations into the chain, when needed.One thing that helps to keep track of the input x as it moves through the chain is using the “.” placeholder (especially when you need to specify function arguments), like so:
f(x) %>%
g(., n=100, param=“baz”) %>%
mean(.$column_name)
Here, the . stands in for “whatever is coming out of the pipe” from the left.dplyr also offers functionality that many don't realize beyond its basic data manipulation. tibbles allow for arbitrary data types for columns, meaning you can have data frames containing your typical strings or floats, but you can also have data frames with columns consisting of nested data frames, or fitted models, or other atypical objects. Once you start working in this way, it can really streamline a lot of complex analytical processes, make your code much cleaner and easier to work with, and allows you to integrate them directly into other tidyverse packages, like ggplot2.
> If I want to use an advanced data manipulation library, I'd typically reach for data.table
Data.table is powerful, but from a usability/readability perspective, most people find it inferior to dplyr. And there are now packages for using the dplyr API and having it run data.table on the backend, so I personally see little to no use for using data.table by itself anymore.
> If I want to use a verb based approach, why not SQL rather than dplyr.
This seems like an overly simplistic reduction. There are bigger differences between dplyr and SQL than both being a verb based approach. And even were there not, dplyr is directly integrated within R. It is still much easier to do your data manipulation directly in R and handing it back and forth between dplyr and whatever other libraries you are using, than it is to do the same in SQL (even utilizing SQL queries within R).
The beauty of the tidyverse is that it is unifying the most important aspects of the data science process under a single approach and API philosophy. It integrates in a way that is pretty unprecedented not just in R, but in any programming language (where data analysis tasks are concerned).
However, there are other tools which predate Tidyverse, which I think so a stellar job at their core competency.
ggplot2 is great, but it was also around long before the Tidyverse concept. But at the same time base R plot() can also be pretty powerful[1] and look great.[2] As an alternative to ggplot2 I would also propose that Vega Lite[3] could be a contender, with an excellent cross language ecosystem.
There are also libraries available for applying SQL on dataframes from directly within R if that's what you want to do. sqldf[4] has been around for a long time, and now there is also the new duckdf[5], which is a bit quicker. Or one can use the DBI[6] library, which requires a bit more coding. Learning SQL is also a great skill which has a lot of value outside R.
Tidyverse may be useful, helpful and convenient for a lot of people, but I think we shouldn't lose sight of the wide R ecosystem which has provided a lot of alternative packages for a long time, perhaps without the marketing and profile of RStudio and the Tidyverse.
[1] http://karolis.koncevicius.lt/posts/r_base_plotting_without_...
[2] https://github.com/KKPMW/basetheme
[3] https://vegawidget.github.io/vlbuildr/
[4] https://cran.r-project.org/web/packages/sqldf/index.html
lattice is largely ignored but it is quite similar to ggplot2 in terms of features, and (as a matter of opinion) the plots it produces are aesthetically more pleasing, also great for making subplots using conditioning variables
My favorite feature of dplyr that makes it stand out compared to SQL is that a window function is just a group_by() without a summarize(). There's no separate syntax for "PARTITION BY".
For data analysis, whenever possible, I don't even write SQL anymore, I use the dbplyr package, which is a dplyr-to-SQL compiler.
* matrices are now a subclass of arrays.
* better reference counting to save on memory.
* stringsAsFactors = FALSE is now default breaking a lot of packages (was this really worth changing?)
* a new way of defining raw strings: r"(...)"
* better default colors for plots. The new default color scheme is less saturated. Apart from looking better, in the old scheme something plotted in bright yellow on a white background was hard to read. Yellow is now more of a mustard.
I know this is popular on the surface, but it's going to break so many "reproducible" analyses from manuscripts and scientific studies.
Furthermore the functions that are inherently about factors (like expand.grid and as.data.frame.table) still use stringsAsFactors=TRUE.
But what are the chances that someone stumbling upon supposedly reproducible analysis happens to know that in R >4.0.0 they need to do this? Especially if the original analysis didn't even specify that it was originally run on R <4.0.0
Definitely worth it. Not sure how many packages it will break and whether it would count as breaking at all.
In my experience, both when I learned R myself back in the day, and teaching others to use it, this is something that trips up newbies over and over again. So from that perspective changing it is good.
OTOH, a lot of existing code will need updating. Oh well..
I’ve never really worked with matrix types in R. What can I do in R4, that I couldn’t in R3? What mistakes will I avoid doing now that matrices are a subclass of arrays?
Also, please don't pay attention to the haters. There are people who hate R, and those who wholeheartedly despise Python; those who loathe Java and plenty of those who can't stand Lisp. Though, R for some reason tends to trigger the most exemplary manifestations of this phenomenon.
Even if they don't agree with the direction Wickham's Tidyverse is going it showcased how flexible the R language was to being rewritten from the inside. Hadley effectively pulled of a more significant language upgrade with less breakages - the R core team could learn something from him here. Even their sensible na.rm default for new functions is introducing more weird inconsistencies to R.
Also they like it this way and promote it: https://twitter.com/hadleywickham/status/1175388442802479104
I like the tidyverse, but Hadley's struggles with lazy evaluation and arguments has cost me lots and lots of time updating internal code at various workplaces.
Don't get me wrong, the tidyverse is great, but if I was writing R code that I expected to run without supervision for a long time, I'd avoid it as much as possible.
Personally I have some tiddy code that is 8 years old and it still is working.
Just like `stringsAsFactors=FALSE` happened on a transition from 3.x to 4.x, because it is breaking, dplyr 1.0 had breaking changes.
Tidyverse seems to be very much aimed at solving the initial learning curve. However, ANY system which uses allows non-experts to fake expertise denies them the ability to achieve real expertise, and also takes away skills from existing experts.
I'm grumpy about some of the reasons given for Tidyverse superiority: The Stringr package is supposedly somehow better because all of its functions start with "string" so you know you are working with strings. Huh? Because it replaces/wraps a lot of functions which begin with grep so you know you are working with regular expressions. Should we also create a "numbr" package which renames every function which works with "numbers"? Or maybe a subset of numbers, maybe we need a package called "integr".
data.table blows tbs out of the water, once you learn how to use it. Yes, it is harder to learn. Yes, it is very much worth it.
Yes, Hadley pulled off a language upgrade, and with the power of RStudio behind him has a strong hand indeed.
Just wonder if it is better or worse?
That is an unusual omission in a current language, and the add on packages are not always a substitute.
Is anyone here aware if that is planned to be a core addition?
The typical programmer's boggled response to R, somewhat like their response to SQL or CSS, is that most of the languages they look at are descended from C, and anything that is not looks weird. Rather like someone who knows only Indo-European languages encountering a non-IE language for the first time.
Of course, any real-life language, whether R or CSS or SQL or C or anything else, has plenty of actual defects to complain about. Like, for example, "=" meaning "change the thing on my left to be the same as what's on my right". But if it's the same defects that you're used to, you don't see them as much.
This confidence seems misplaced. Experienced developers using better languages with better tooling create bugs all the time. How likely is it that R packages, written by specialists (or grad students) whose main focus is often not programming but some other discipline, in a language full of "quirks," are going to be really reliable?
One quirk that got me the last time I touched R was that functions and variables live in different namespaces, so you could assign 4 to foo and then call it.
Fun fact, a lot of the R source code is machine translated from Fortran to C, but it has been cleaned up a lot. https://github.com/wch/r-source/search?p=1&q=f2c&unscoped_q=...
Either I'm greatly misunderstanding or this is just plain wrong - assigning something to the same name as an existing function will just overwrite the function.
my_vec <- c(1, 2, 3)
c <- 4
print(c)
my_vec2 <- c(5, 6, 7)
But as I describe in my top level comment, it has nothing to do with variables versus functions, just how namespaces work. z <- function() { print("hello world") }
z <- 3
z # Outputs 3
z() # Errors because there is no function z()
What the actual quirk you're talking about has to do with attaching other namespaces. So, the way R works is that by default it loads several packages. For instance, the reason you can call `mean()` is that the "base" package is attached when you open a blank session in R. It is absolutely possible to have multiple symbols take the same label across namespaces. Here's an example: mean <- 4 # Creates variable called mean in the local namespace
mean(c(5, 10, 20)) # Uses the mean function in the attached base namespace
mean # Uses the mean variable in the local namespace
But this is not unique to variables versus functions, it also works with functions: mean <- function(x) { sum(x) * 100 }
mean(c(1, 2, 3)) # Outputs 600
base::mean(c(1, 2, 3)) # Outputs 2
rm(mean) # Drop local namespace function
mean(c(1, 2, 3)) # Outputs 2
You can use sessionInfo() to see the order of attached namespaces that will be searched. In RStudio, you can also press the little dropdown arrow next to "Global Environment" in the environment pane to see the order of attached namespaces -- it'll search them in that order. Alternatively, you can be defensive and always prefix all functions in any namespace with the namespace name (equivalent to always using module.function style function calls in python).Finally, you can use the built-in function "find" to see the order in which R will try to resolve a symbol, e.g.
sum <- function() { print("i like sum coding in R!!!") }
find("sum") # .GlobalEnv first, base second
I'm not sure this is a quirk. What are your options as a namespaced language? 1) Allow users to import multiple functions/variables with the same name across different namespaces and resolve the conflict via some kind of hierarchical order; 2) Don't allow users to import multiple functions/variables with the same name, so you can never use overloading to monkey patch; 3) Always require namespace prefixes at all times; 4) Make function dispatch a blocking operation that asks REPL users which to use? I dunno, it doesn't seem to me like R's approach is any less sane.I guess one thing that is different about R's approach is that "built-ins" have no special priority, they're all part of some namespaces that are attached by default but otherwise exactly like third-party libraries (base, stats, graphics, etc.)
In R, the only reserved words that cannot be overloaded are while, repeat, next, if, function, for, and break. (Note: else does not appear here because of a genuinely baffling quirk about how else is implemented in R)
Actually, to answer my own question (I know google), I found this very informative stackoverflow answer [1]. My TLDR: no difference, other than "<-" is more likely to cause carpal tunnel syndrome.
[1] https://stackoverflow.com/questions/1741820/what-are-the-dif...
What I find distasteful is that when calling mean(), the resolution of this name depends NOT on whether the local variable mean has been defined, but whether it has been assigned a function. This is illustrated by your 2nd and 3rd examples.
Of course if you are used to it, it may not catch you by surprise.
For example:
mean(2, 4) # outputs 2
mean(c(2, 4)) # outputs 3
I really don't see why R allows you to enter mean(2,4) without giving an error.One solution I have found for small and medium data is to use rpy in Jupyter to let me keep most of my workflow in python, then shuttle stuff to R for exotic tests or to use key packages (ggplot, brms, lme4).
It's comparatively very weak for "software development" - writing libraries, long-running programs, etc. Even data cleaning/data analysis scripts are a headache beyond a certain length or number of contributors. People tend to have the same complaints about MATLAB, Stata, and Mathematica.
These are related - a lot of the features that make interactive use concise and easy to pick up end up being kind of a mess to handle consistently in a larger program.
Python is at a different point on this spectrum. It's moderately good for interactive analysis, but there are still a lot of tasks that are concise and easy in R that require a bunch of verbose object manipulation in Python. This is a choice - "explicit is better than implicit", etc. Python is also moderately good for "software development" - it has enough consistency and code structuring features that you can write libraries, systems, and infrastructure while collaborating with a small number of other people pretty easily. You still hit a point where Python's dynamic features make things like refactoring harder than you'd like, but it's usable.
I used to do mostly data analysis in my day-to-day work and R was my go-to and absolute favorite language for years in terms of usability for data analysis. Doing actual software development in R is quirky at best, to be honest.
Nowadays I write code for research that requires 'actual' software development, so I've been using python almost exclusively (with pytorch under the hood, which I love.) No doubt, python is a better language for software engineering.
Nevertheless, for analysis I just cannot warm up to numpy/pandas/matplotlib _at all_. When it's time to analyze results of my experiments or produce publication level graphics, I write my python results to disk and use the tidyverse as a last mile solution.
The most famous language with this property is Lisp.
We can have a $foo variable, and a foo command/function.
GNU Make is another one, sort of.
Makefile:
warning = abc
$(info abc = $(warning))
$(warning what?)
Output: abc = abc
Makefile:3: what?
make: *** No targets. Stop.
$(warning) is a variable, whereas $(warning args ...) is an operator call.If a macro is stored in a variable V, it cannot be called as $(V args), but using $(call V args). That's analogous to funcall in Common Lisp.
Looks like the Lisp-2 approach is well represented in the famous language scene. :)
The two namespace approach makes it harder to assign 4 to foo and then call it, so I don't understand this comment.
You can assign 4 to foo and try to call it in a language that has one namespace.
;; Scheme, one namespace:
(let ((foo 4))
(foo) ;; oops
;; Common Lisp, two spaces:
(let ((foo 4))
(foo) ;; still calls the foo function,
;; not related to or shadowed by the above variable.
(funcall foo)) ;; oops
The usual valid complaint about two namespaces is that funcall is required all over the place in code that works with higher-order functions, which uglifies the code, and that an operator like (function foo) or its abbreviation #'foo is required to lift references to functions as values instead of just writing foo, which likewise uglifies code.This is clownshoes software quality. And it's used for important scientific research.
Do you have the data to back that up or is just a question that you don't have an answer to?
I think the statement brush off the high quality R packages that already in the ecosystem that is not in anywhere else. The author of element of statistic learning created glmnet package which took python years to even have ported. There are many other packages out there that Python does not have.
I am not going to argue it's the prettiest language.
But there are tons of packages and many found no where else in other language ecosystem.
If you're going to say just use rest/rpc then it just defeat the purpose of your initial argument.
And just look at springer or, gosh the other publisher escaped my mind right now, but these publishers have book on statistical subjects with R packages accompany the book.
It may be ugly but the packages are maintain by expert in the field of statistic and they may not be programmer. But they do dog food their package and use dataset to see the results. Likewise just read up on the Ranger Rpackage paper (https://arxiv.org/pdf/1508.04409.pdf). They test their output with the other randomforest package.
And it's silly to point this argument at just R when the same happen to Python. The bootstrap function in SciKit-Learn for the longest time didn't even really do bootstrap. The linear regression function automatically does shrinkage with no option to turn off.
No language is perfect. But I believe R have a place and especially in statistic. Many many wonderful expert statistician are maintain and creating R packages (eg Dr. Frank E. Harrell Jr. ) Many R packages have accompanying paper publish here https://www.jstatsoft.org/index
The thing about a lot of R code not being written by programming specialists is a valid point, but then again not many programming specialists are also specialist in statistics so... the alternative is what?
I tried to do a simple word cloud based on a Twitter search in R. The whole thing was plagued by mysterious failures due, AFAICT, to weird, rare, non-standard UTF characters in the source data that were crashing some of the R libraries I was using to clean the data.
Having done that before in Python, it was shocking to see how fragile the R ecosystem is.
Gotta say, though, one of the most annoying things about R is the name. Entering “R” as a search term on job boards tends to lead to a lot of not-very-helpful results.
Also, it’d help if the version release names had more order to them. So you’d read some ridiculous phrase like “parachute trombone” and know that it must’ve been released after “mouse parade”.
It's a new major version - does it require significant updates for training materials? Are there a lot of outdated idioms now?
If you're regularly running into memory performance issues despite using data.table, and you don't feel like fiddling with mmap, consider looking into an APL variant.
Also pull important stuff to the top of announcement. Were there security issues that need you to upgrade? What are the major new features that would encourage you to upgrade?
For example here's a release announcement I wrote recently: https://www.redhat.com/archives/libguestfs/2020-February/msg...
Free software doesn't usually have an advertising budget, so you have educate people on what your software does at every opportunity you get.
I think it's reasonable to assume that this message was intended for subscribers of that mailing list, and that the people who chose to subscribe already know what R is.
Request to please start release announcements by saying what really a "Network Block Device (NBD) server" is. Don't assume that people are familiar with every possible piece of software already.
Of course there's a level beyond which you don't really need to go - I wouldn't suggest explaining what Linux is or what software is.
Release notes are usually for people who are already using it.