R beats Python, R beats Julia, Anyone else wanna challenge R? (2014)
matloff.wordpress.com
matloff.wordpress.com
You missed the point. Everything is fast when you export what you're doing to a library. Julia is fast when you do, and when you DON'T.
P.S. that function name though...
In modern times, vector based thinking like that that's present in R or numpy is the only sane way to program. Because these are not "library calls," they are the primitives of modern architectures.
I think that scientists are much more likely to enjoy the lisp nature of R than the code-monkeys that get caught up because arrays start with 1 instead of 0.
That confused me a great deal. Julia seems more lispy than R. From the comment section it seems that you have perhaps implicitly decided apriori that R is better. So you would rationalize out of evidence and examples of situations to the contrary.
My interpretation of the term "data science" is much more in line with practicality of engineering rather than the purity of traditional statistics. To that end, Python is leagues ahead.
He's right. 'Machine learning' is just statistics with better, cleaner names for things. The underlying theory is the same.
(The commenter below says that 'statistics explains, ML predicts', but this isn't really true. Both statistics and ML build models, whether you use these models for prediction or explanation is up to you.)
One problem is that the 'statistics' as usually taught in a college course or explained in a textbook comes from a much earlier age, before supercomputers were available to the average man; in essence, it's like a kind of machine learning as done by paper-and-pencil. In contrast, 'ML' assumes from the start that computing resources will be available.
Right, but I'm not sure that "better, cleaner names for things" actually follows. Instead, I find that the ML folks just hacked their way to similar results as traditional statistics, but in many cases were comfortable with the algorithms as "black boxes" rather than having a clear understanding of why the algorithms worked. In that sense, the author's "unprincipled" criticism is valid. This is less true today, but the new research in convolutional neural nets shows how ML starts from hacking things until they produce practical results then backing into the theory of why. This habit has resulted in much duplication of effort and naming schemes. My ML prof at Georgia Tech (Isbell was awesome! http://www.cc.gatech.edu/~isbell/) constantly trashed "genetic algorithms" for being a silly form of randomized hill-climbing.
The beneficial side of these less-principled techniques is that they happen to work on larger scale datasets. It turns out approximate results are more scalable than exact results.
"This has been the subject of huge controversy over the years, so Guido Van Rossum, inventor of the language, added a multiprocessing module."
I think that assigns too much agency to van Rossum. PEP 371 describes the justification for bringing pyProcessing into the standard library. The primary developer was Richard Oudkerk, who along with Jesse Noller volunteered to maintain it in the standard library.
While it's true that van Rossum accepted the PEP (on Thu, Jun 5, 2008 at 1:22 PM according to the mailing list), that's not quite the same thing as adding it.
For that matter, the version control logs point out who originally added the code, in its most literal sense:
user: Benjamin Peterson <benjamin@python.org>
date: Wed Jun 11 02:40:25 2008 +0000
summary: add the multiprocessing package to fulfill PEP 371As both a heavy R user and a software engineer, I can promise you that one of the quintessential aspects of R is it's actually "written by statisticians, for statisticians."
You can't accuse R of having great code and language design. Or good code and language design. Or even mediocre code and language design.
Imagine what you would get if you got a million monkeys drunk, put them on a roller coaster with laptops, and had them bang keys while they were upside down on loops. And then the result suddenly, miraculously runs and produces output. Now you understand R's software design.
cat(paste("That", "and", "string", "manipulation", "in", "R", "is", "a", "pain", "in", "the", "ass", sep=" "))
[1]: If A is n by m and B is p by q, and m is a multiple of p, then R will SILENTLY concatenate copies of B to itself to form an m by q matrix, and then do the multiplication.Example, please. I'd really like to see what you are talking about here.
And btw, good luck treating any kind of Big Data with JMP.
Oh I'm definitely not intending to. (To the commenter below, it really doesn't matter how you define Big Data for that statement.) That much I had already figured out, which was why I was evaluating R with an eye towards pitching my boss on it.
The test case I'm referring to here was a pretty simple neural net to my mind -- roughly 350,000 rows of data, six predictor variables, one hidden layer with 30 nodes. I can verify that the neural net code ran because if I truncated the data set down to 1,000 rows I got a result back. But the full dataset just chugged for hours and hours without stop.
R has some nice individual packages, and some bright people involved, but the core is rotten. While it is possible to write high performance R, and possible to write easy-to-read R, I don't think it's possible to do both at the same time. It's easy to drop to C or C++ to speed up the critical sections, but often the glue between the pieces is so slow that it's not worth bothering.
I'm not recommending that you use Perl for statistics (Python or Julia would be better choices), but I think a language written by a linguist for sysadmins is a much better choice than a language written by statisticians for statisticians. Although maybe that's because I'm a programmer and not a statistician?
Some of my sampling code was running much slower than I thought made sense (even for R), and the strange part was that larger samples were sometimes 1000x faster than smaller samples. The answer was that R (wisely) uses two algorithms for sampling without replacement, but the cutoff between them is based on a fixed sample size of 1e7, rather than the percentage of population sampled. Perhaps there's a good reason?
Case in point:
- stringsAsFactors is True by default, leading to all kinds of weird silent behaviors
- length('string') being 1, since that's silently a vector of length 1 (~ length(c('string')) ). Have to use nchar() or str_length()
But, analyzing strings is pretty cool - stringdist() or levenshteinDist() are great!
Data frames in Pandas have no such implicit stupidity. The rough equivalent, IIRC, is setting up one or more columns as an index (or multi-index, respectively), and that's mostly done explicitly (with the exception of importing or exporting CSVs; the first column is imported as the index by default, and the index values are written out by default).
Oh, and you think all programmers from all time were all born with the same background and same experience and CS degree?
Yeah, like the Wright Brothers were aviators and aircraft designers before they made their own plane ?
Well I think that everybody will agree that R is not good software. It's still useful, though.
A couple of idiotic things about R have already been posted in this thread, let me add one of my pet peeves: "how does one get the path of the currently executing script?". Last time I looked, the answer was "there is no way to do so without relying on implementation details that can and have changed between minor releases.". Lol, that's pathetic.
Now that Microsoft has made substantial investments in R, I hope that at least they'll fix the major issues in embedding R through the C api. (and no, Rcpp is substandard - it doesn't even support the msvc compiler).
What are we supposed to glean from this paragraph?
1. Not good to make assertions that you immediately undermine by admitting you don't know what you are talking about 2. I find @spawn and @fetch quite nice, in fact I think that Julia was built from the ground up around notions of parallelism (ie. multiple dispatch)
> There is a huge discussion about this on the mailing list; please see that. If 0 is mathematically "better", then why does the field of mathematics itself start indexes at 1? We have chosen 1 to be more similar to existing math software.
A post above also appeals to mathematica. It's been a while since I've had to internalize the "cost" of zero indexing, so this never felt very compelling.But it's true, indexing from 1 is common, or idiomatic, in Fortran.