As both a heavy R user and a software engineer, I can promise you that one of the quintessential aspects of R is it's actually "written by statisticians, for statisticians."
You can't accuse R of having great code and language design. Or good code and language design. Or even mediocre code and language design.
Imagine what you would get if you got a million monkeys drunk, put them on a roller coaster with laptops, and had them bang keys while they were upside down on loops. And then the result suddenly, miraculously runs and produces output. Now you understand R's software design.
cat(paste("That", "and", "string", "manipulation", "in", "R", "is", "a", "pain", "in", "the", "ass", sep=" "))
[1]: If A is n by m and B is p by q, and m is a multiple of p, then R will SILENTLY concatenate copies of B to itself to form an m by q matrix, and then do the multiplication.R has some nice individual packages, and some bright people involved, but the core is rotten. While it is possible to write high performance R, and possible to write easy-to-read R, I don't think it's possible to do both at the same time. It's easy to drop to C or C++ to speed up the critical sections, but often the glue between the pieces is so slow that it's not worth bothering.
I'm not recommending that you use Perl for statistics (Python or Julia would be better choices), but I think a language written by a linguist for sysadmins is a much better choice than a language written by statisticians for statisticians. Although maybe that's because I'm a programmer and not a statistician?
Some of my sampling code was running much slower than I thought made sense (even for R), and the strange part was that larger samples were sometimes 1000x faster than smaller samples. The answer was that R (wisely) uses two algorithms for sampling without replacement, but the cutoff between them is based on a fixed sample size of 1e7, rather than the percentage of population sampled. Perhaps there's a good reason?
Example, please. I'd really like to see what you are talking about here.
And btw, good luck treating any kind of Big Data with JMP.
Oh I'm definitely not intending to. (To the commenter below, it really doesn't matter how you define Big Data for that statement.) That much I had already figured out, which was why I was evaluating R with an eye towards pitching my boss on it.
The test case I'm referring to here was a pretty simple neural net to my mind -- roughly 350,000 rows of data, six predictor variables, one hidden layer with 30 nodes. I can verify that the neural net code ran because if I truncated the data set down to 1,000 rows I got a result back. But the full dataset just chugged for hours and hours without stop.
Case in point:
- stringsAsFactors is True by default, leading to all kinds of weird silent behaviors
- length('string') being 1, since that's silently a vector of length 1 (~ length(c('string')) ). Have to use nchar() or str_length()
But, analyzing strings is pretty cool - stringdist() or levenshteinDist() are great!
Data frames in Pandas have no such implicit stupidity. The rough equivalent, IIRC, is setting up one or more columns as an index (or multi-index, respectively), and that's mostly done explicitly (with the exception of importing or exporting CSVs; the first column is imported as the index by default, and the index values are written out by default).
Well I think that everybody will agree that R is not good software. It's still useful, though.
A couple of idiotic things about R have already been posted in this thread, let me add one of my pet peeves: "how does one get the path of the currently executing script?". Last time I looked, the answer was "there is no way to do so without relying on implementation details that can and have changed between minor releases.". Lol, that's pathetic.
Now that Microsoft has made substantial investments in R, I hope that at least they'll fix the major issues in embedding R through the C api. (and no, Rcpp is substandard - it doesn't even support the msvc compiler).
Oh, and you think all programmers from all time were all born with the same background and same experience and CS degree?
Yeah, like the Wright Brothers were aviators and aircraft designers before they made their own plane ?