Comparison – R vs. Python: head to head data analysis
dataquest.io
dataquest.io
But when you start having to massage the data in the language (database lookups, integrating datasets, more complicated logic), Python is the better "general-purpose" language. It is a pretty steep learning curve to grok the R internal data representations and how things work.
The better part of this comparison, in my opinion, is how to perform similar tasks in each language. It would be more beneficial to have a comparison of here is where Python/Pandas is good, here is where R is better, and how to switch between them. Another way of saying this is figuring out when something is too hard in R and it's time to flip to Python for a while...
Out of curiosity, why do you consider CRAN to be much better than PyPI?
And the last time I checked, "pip install numpy" could be quite a pain, especially if you needed to compile dependencies. Rstudio makes it ridiculously easy to install R and add packages.
However - for all other types of packages, PyPI is obviously superior. The breadth of packages on PyPI is much better than CRAN.
It about choosing the right tool for the job.
To be honest, matplotlib seems a good contender to me (http://matplotlib.org/).
Also, what's wrong with comparing R to Pandas/Numpy ? They can only be used from within Python, right?
Edit: just realised from another comment that Pandas/Numpy can be accessed from R, too.
Absolutely nothing.
I was referring to the article title that is was an R vs. Python comparison. Python is so much more in terms of a general purpose language than R is. Similarly, R is much more in terms of stats (built-in) than Python. I just thought that it would be more accurate to call the article an R vs. Pandas/NumPy comparison.
Even though both of them need an extra plotting library to make publication quality plots. Matplotlib isn't bad by any means - and it's gotten better over the years. But R/ggplot2 produces nicer plots (IMO). I'm not sure that I'd export data from Python into R just for ggplot, but I might.
I am not that familiar with ggplot myself, but I'll give it a go as soon as I'll have the chance.
> To be honest, matplotlib seems a good contender to me (http://matplotlib.org/).
They're quite different, though, and I can see why many prefer ggplot. It's a declarative, domain-specific language that implements a Tufte-inspired "grammar of graphics" (hence the gg- in the name; see section 1.3 of [1], and [2,3]) for very fast and convenient interactive plotting, whereas matplotlib is just a clone of MATLIB's procedural plotting API.
[1] http://www.amazon.com/ggplot2-Elegant-Graphics-Data-Analysis...
[2] http://www.amazon.com/The-Grammar-Graphics-Statistics-Comput...
On paper perhaps, less so in application. Sure you can probably make matplotlib do everything ggplot does with enough work, but working with ggplot is just so much quicker easier and more fun.
And I say that as someone who does all his data analysis in Python.
You can try out my dev version [1] (rewrite branch). It will be nearly API compatible.
I've waxed lyrical about Python all over this thread, but here you have to give the medal to R. Matplotlib is one of my least favourite libraries to use, been doing it for almost 2 years, and I still spend half my time buried in the documentation trying to figure out how I'm supposed to move the legend slightly to the right or whatever.
ggplot probably has slightly less flexibility overall (mpl is monolithic), but for just doing easy things that you need 99% of the time, ggplot is king.
I am not familiar with ggplot, so I wasn't comparing them on the ground of the easiness of use, but by looking at some ggplot examples, they looked like something you can do with matplotlib, too, so I pointed that option out, too.
On the other hand, with ggplot, you can create a good enough chart in couple of lines of code almost for any data.
BTW, there's a ggplot port for a Python: http://ggplot.yhathq.com/
I tried the ggplot for python (ggplot.yhathq.com/) but eventually settled for seaborn (http://stanford.edu/~mwaskom/software/seaborn/). It is really quite easy to get most of the common plots that I wanted and hasn't let me down yet. The standard plots look SO much better than the standard plots of MPL without a lot of customization.
And yet, you go right on in the next sentence to make it a Python/Pandas/Numpy vs. R/everything in CRAN comparison. Libraries count.
you can code in multiple languages in your noteobook, and they can all communicate, making it easy to go from Python to R to JavaScript, seamlessly.
we just released v1.4 with all kind of new features, check it out: https://github.com/twosigma/beaker-notebook/releases/tag/1.4...
I didn't get it working on my Linux machine, but you will definitely see some pull requests once I have time to fiddle with it. The electron version is a nice idea but I would prefer better instructions for installing the normal version. "This script will do it all" is not always helpful.
We are working on better linux packing and distribution (see our issue tracker), but it is not easy to do it right, and it will take a while.
PRs very welcome!
Matplotlib with some good definitions ends up providing much better results and nicer looking plots fro B&W unlike what people normally think.
Yes; Python is a better general purpose language. It is inferior though when it comes specifically to statistical analysis. Personally I don't even try to use R as a general purpose language. I use it for data processing, statistics, and static visualizations. If I want dynamic visualizations I process in R then typically do a hand off to JavaScript and use D3.
Another clear advantage of R is that it is embedded into so many other tools. Ruby, C++, Java, Postgres, SQL Server (2016); I'm sure there are others.
My 2 cents: If someone has no programming background, then building a foundation from python will allow them to do much much more than building a foundation on R--unless of course they only care about statistical analysis and have no inclination to code more generally. I learned both at the same time even though I had no use for Python at the time (was and still am a professor) but I use it almost everyday now and very much enjoy it!
I'd say R is a _terrible_ language. Its types are just really different from every major programming language, and it's horrible for an experienced programmer to use.
I totally agree that R has fantastic libraries, but I'd like to see people focus on improving libraries for Python rather than sticking with R, which as a language is less well-designed than Python.
[I use R for most of my stats, I also use Matlab and Python]
Without any standards at all, people would have at least produced a readme.txt, which would have been a huge improvement -- e.g. I much prefer working with unfamiliar user-written Matlab packages :)
https://stat.ethz.ch/R-manual/R-devel/doc/html/packages.html
library('zoo');
vignette('zoo');
#####
library('ggplot2');
vignette(package='ggplot2')
My experience was exactly the opposite -- first time I saw R syntax (actually, it was S-Plus back then...) , I thought it was the most intuitive and powerful system I've ever seen -- this was after fairly extensive experience in C and C++, as well as a few others.
Now, I don't quite think so any more, because there are many rather tricky things buried under the surface (e.g. how many people really understand how exactly environments work?) -- but the majority of R programmers will never have to deal with them in their code...
Also, I have definitely done general-purpose coding in R -- for a lot of things it is completely adequate. Python has more general-purpose functions and libraries of course, similarly to how R has more statistical ones.
That said, the language is often taught poorly. Here's my attempt to do better: http://adv-r.had.co.nz
Thank you for all of your hard work! Keep on keeping on; your contributions have been phenomenal!
That sounds promising, I'll check it out, thanks.
I think R is a great tool, but I maintain that it is not a well-designed language by modern standards.
However, from a computer language design point of view - it leaves a lot to be desired. It's type system is seems very complicated and while the language tries to do what it thinks you want, it's not always clear what is going on (are you working on a matrix or a dataframe that has been cast into a matrix?).
For me, R is one of those languages that is good in a certain domain, but once you get out of that domain, it makes things more complicated than they need to be. It just isn't a general purpose language. By far, the biggest problems I've seen have been people who only know R (mainly stats people or biologists) try to do something in R that would be a quick 10 line Python/Perl/Ruby/whatever script.
Normally for a language design, you aim to make easy things easy, and difficult things possible. For R, it seems like it makes difficult things easy and easy things difficult. Maybe that's the tradeoff that was needed. :)
That said - please keep doing what you're doing. You've made my R work vastly easier.
- http://stackoverflow.com/questions/1815606/rscript-determine-path-of-the-executing-script
- http://stackoverflow.com/questions/3452086/getting-path-of-an-r-script
(where you already commented, so it's not like this is something new...)I would say that any language that does not have a facility to get the path of the current file, is not 'excellent' under the criteria an experienced programmer would use for assessing it.
Now, I very well know that those criteria are different from what scientists use, but still...
I have to disagree. Its main model is generic function method dispatching. It can feel odd at first to someone coming from the C++ style of OO where objects own methods, not methods owning objects. But it's a legitimate OO style with its own advantages. [1]
I've found the more I use R, the more intuitive a lot of its operations are. It's relatively easy to "guess" what you ought to do to accomplish what you want. More so then other languages I've learned.
It's not the worst language in the world, but it isn't terrific language either.
I like R as a higher level language (or I guess tools like SPSS or preferably PSPP for even higher level stuff). These days I do most of my academia stuff with R (mostly hypothesis and equivalence testing and the things related to it like power analysis etc.)
I've never really looked into Python which is strange because I use it as a "glue language" quite often. I think I'll investigate Python a bit more next time I have to actually collect and clean up the data before using it. Right now I'm more of a consumer (mostly using data from our experiments that are turned into CSV)
It is indeed. And R works with Fortran quite easily.
> I like R as a higher level language (or I guess tools like SPSS or preferably PSPP for even higher level stuff). These days I do most of my academia stuff with R (mostly hypothesis and equivalence testing and the things related to it like power analysis etc.)
You can see R as some sort of glue language around libraries written in lower languages like C++, C or Fortran (I believe a large part if not all the functionalities for matrix operations used by R for linear regressions and statistican analysis (PCA) is written in Fortran).
Fortran code runs much faster, but you don't want to use it to do exploratory analysis ("I have those data about people, what if I filter out the people earning more than X before checking if there is a correlation between the average age where men get married and their incomes?").
The main difficulty with Fortran is IMO the lack of an extensive standard library -- sure, you can find code out there to do almost anything, but then you need to figure out linking/calling conventions/possibly incompatible data models for each new library you bring in...
But, as another poster mentioned, it is quite straightforward to call Fortran from R :)
It's not like R does not have obtuse and baroque parts, it certainly does, and their obtus-ity is rather high, but IMO they are not parts of the language a casual user would likely encounter...
On the other hand, Python has quite a few pitfalls itself -- but I suspect a casual user would, for example, run into Python default arguments a bit sooner than she would run into R environments :)
https://cran.rstudio.com/web/packages/dplyr/vignettes/introd...
End example from a airplane arrival and departure dataset:
flights %>%
group_by(year, month, day) %>%
select(arr_delay, dep_delay) %>%
summarise(
arr = mean(arr_delay, na.rm = TRUE),
dep = mean(dep_delay, na.rm = TRUE)
) %>%
filter(arr > 30 | dep > 30)There's scant few articles on going from Python to R...and I think that has given me a lot of reason to hesitate. One of the big assets of R is Hadley Wickham...the amount and variety of work he has contributed is prodigious (not just ggplot2, but everything from data cleaning, web scraping, dev tools, time-handling a la moment.js, and books). But that's not just evidence of how generous and talented Wickham is, but how relatively little dev support there is in R. If something breaks in ggplot2 -- or any of the many libraries he's involved in, he's often the one to respond to the ticket. He's only one person. There are many talented developers in R but it's not quite a deep open-source ecosystem and community yet.
Also word-of-warning: ggplot2 (as of 2014[1]) is in maintenance mode and Wickham is focused on ggvis, which will be a web visualization library. I don't know if there has been much talk about non-Hadley-Wickham people taking over ggplot2 and expanding it...it seems more that people are content to follow him into ggvis, even though a static viz library is still very valuable.
[1] https://groups.google.com/forum/#!topic/ggplot2/SSxt8B8QLfo/...
Also worth pointing out, he's actively working on a new book for ggplot2, which, AFAICT, he's providing for free (you just have to run the build tools)
https://github.com/hadley/ggplot2-book
I think if someone were to run an analysis of Wickham's Github activity, it would produce a freakishly busy chart.
I used to work a lot with R many years ago. I was shocked to find how bad the documentation was, and worse how rude and unfriendly the "community" of grumpy professors was. I shudder to think of the horrible meanness towards beginners asking questions on the mailing list.
I got so fed up I even wrote a book about R data visualisation. But this was all just around the time ggplot2 came out. Unfortunately I stopped using R soon after, but since then Hadley has single-handedly done more good for the language than anyone else.
I don't know what the R community is like now, and whether people like Hadley have made it friendlier, but it's clearly one reason Python is superior.
On the other hand, there seem to be a lot of useful libraries that haven't been ported over to Github or are otherwise easily accessible beyond CRAN...Many of them probably don't get as much exposure as they would if they were more easily discoverable...and I honestly don't even know where, in those cases, to start the bug reporting/patching process. That's obviously the fault of my being spoiled by Github...but that's kind of the point, there's a bit more friction in contributing to R than you might find in Python/Ruby/etc.
The caveat on the ggplot2 book is that building it seems to be really hard because of the nightmare of cross-platform latex. But there will be a physical book out early next year.
Every language has third party packages that are primarily the work of one person.
I'm sure your statement is true for some definition of deep but I don't agree.
This isn't to say that there aren't other programmers doing brilliant work in R (also, R is just a smaller community overall), but he's devoting significant time to building out support tools and frameworks...this suggests that he is a total mensch, but also that there was a significant need that hadn't yet been addressed.
Another interpretation is that R is an incredibly productive language for this sort of programming, otherwise one person couldn't write so much useful code. ;)
However when the dataset is medium sized (i.e.: fits into your computer's memory / 2) R crushes Python (and Pandas) for the 80% of the time you'll be spending wrangling. The reason is that R is vector-based from the ground up. Pandas does everything that R does, but does it in a less-consistent, grafted-on way, whereas the experienced R person who "thinks vectors" is way ahead of the Python guy before the analysis has even started (i.e., most of the work). I know both really well. I use Python when I want to "get (semi) serious" production wise (I qualify with "semi" because if you're really serious about production, you're probably going to go to Scala).
But when it comes to taking a big chunk of untidy data and bashing it around till it's clean and cube-shaped, will parse, and has no no obvious errors, R is miles ahead of Python. R is where you do your discovering. Python can do it too, but I would estimate the cognitive overhead as double.
By the way, that's why people who "think time series" all day long (i.e., vectors, not objects), and who want to implement their algos, not think CS, will first typically build it in R, which is why CRAN beats Python all the time and every time for off-the-shelf data analysis packages. Data people go to R, computer-people go to Python (schematizing).
R is slow. That's its main problem. And that's saying something when comparing it to Python! But the gem of vector-everything makes it a much more satisfying language than imperative, OO, Python, when it comes to the world of data first, code second.
Finally I'd add that Python 3.x is arguably distancing itself from the pragmatism which data science requires, and 2.x provided, towards a world of CS purity. It's not moving in a direction which is data science friendly. It's moving towards a world of competition with Golang and Javascript, and Java itself.
http://www.johnmyleswhite.com/notebook/2013/12/22/the-relati...
had worried me a couple of years ago. JMW shows that vectorized was much slower also in Julia (though still both faster than R - but that's not difficult).
Glad to see Julia is very fast in both cases, though it's still somewhat perplexing the extent to which vectorized code is necessarily slower. I'm thinking that the future of GPU enabled languages will mean vectorized code will be faster, so I prefer languages with a bias towards vectorisation.
The vectorized code typically allocates all kinds of intermediate results (more GC, more memory accesses). Apparently, turning it into loops is less trivial than it seems.
I'm thinking that the future of GPU enabled languages will mean vectorized code will be faster, so I prefer languages with a bias towards vectorisation.
I share that concern. Julia has some libraries to support GPU programming, but I don't know of any plans to have the core compiler take advantage of it.
However, devectorization (i.e. replacing vector ops with a for-loop) is sometimes a performance improvement because Julia can usually provide C-like speeds in for-loops and avoid creating intermediate arrays.
[1] http://www.johnmyleswhite.com/notebook/2013/12/22/the-relati...
[2] http://blog.rawrjustin.com/blog/2014/03/18/julia-vs-python-m...
I am a data person, and I have to deal with a lot of text in my job. If I had to do it in R, I would quit.
Can you explain why you think it is easier to wrangle data in R? My experience is the opposite.
Perhaps I should clarify, I'm talking mainly time series and/or data which is vectorizable. Python is better if you're scraping the web. If there's a lot of if else going on. Ie imperative programming.
R's native functional aspects (all the apply family) and multilevel vector/matrix hierarchical indexing is better built from the ground up for large wrangling of multivariate datasets, in my opinion.
> rollapply(some1000x10matrix, 200, function(x) eigen(cov(x))$values[1], by.column = FALSE) # get the first eigenvalue rolling 200x10 window.
>>> # impossible in Python unless using ultra-complex Numpy stride tricks.
> dim(someMatrix)
>>> someMatrix.shape
> head(someMatrix)
>>> someMatrix.head() # notice consistent function application in R, whereas in Python, mixed attribute / function? So we're on OO land and I must know if it's an attribute or a function....
> rollapply(some1000x2matrix, 200, function(x) {linmod <- lm(x[, 1] ~ x[, 2]); last(linmod$residuals) / sd(linmod$residuals)}, by.column = FALSE) # get the z score in one multi-step function.
>>> Impossible in python without For loop as lambdas cannot be multi-statement.
> native indexing using [] brackets by index number, or index value, or boolean. All vectors.
>>> pandas loc/iloc/ix mess.
> ordered lists (python dict) by default, so boolean or index subsection easy even when data is hierarchical, not tabular
>>> easy bugs due to unordered nature of dicts; must import some different module and then still can't vector index it.
It's all summed up by this: > c(1, 2, 3) * 3
[1] 3 6 9
>>> [1, 2, 3] * 3
[1, 2, 3, 1, 2, 3, 1, 2, 3] # wrong! Need rescuing by Numpy!
And then there's CRAN. Just last night someone told me about "nowcasting" which uses "MIDAS regression". A relatively new technique. Google it for R (full package available), Google it for Python (Matlab comes up ;-).And I'm not even going to start on graphics. Seaborn and bokeh are valiant efforts, but they're still 80% of what ggplot and base graphics can do, especially, at the multidimensional scale. That last 20% is often all the difference between meh and wow. That said, I do appreciate Matplotlib's autos rescaling of axes when adding data. Python charts aren't as pretty nor capable of complexity (for similar effort), but they're arguably more dynamic.
Now don't get me wrong. The converse list for Python would be much longer, because it's more general purpose, and it kills R outside of data science. I wrote 10k loc in R for a semi-production and it was horrible because it does not have the CS tools for managing code complexity, and it really is slow at certain things. R is more focused on iterative, exploratory data science, where it excels.
Sorry, couldn't resist!
On the other hand, if you know the exact calculations that you need to do and the results you're gonna get, then Python might be a better tool.
Personally I learned R after Python, and I use both languages, but I prefer R for anything involving statistics.
Have you tried IPython/Jupyter?
What I meant by better interactive interface is that the language itself is designed with interactive use in mind.
For instance compare
func(x$a, x$b)
func(a, b, data=x)
func(a=1, b=2)
with func(x['a'], x['b'])
func('a', 'b', data=x)
func([1, 2], ['a', 'b'])
The R versions are easier to type and read.I think tidyr also currently has the edge over pandas for making [tidy data](http://vita.had.co.nz/papers/tidy-data.html).
Also I assume that by "something non-standard" you mean something other than a way to analyze it? Because there is really no comparison wrt available analysis packages between the two...
Not trying to say that R is perfect and great for everything, definitely not, I just have a hard time imagining a data-processing task for which I would choose Python over R (I might pick SAS over either one of them though...)
While there is (almost?) always a way to do a SQL query using idiomatic R, I have to admit that sometimes my brain thinks up a solution in SQL faster (a product of upbringing).
You've got R Studio, which is one of the best environments ever for exploring data, visualisation, and it manages all your R packages, projects, and version control effortlessly.
Then you've got the plethora of packages - if you're any of the following fields: statistics, finance, economics, bioinformatics, and probably a few others, there's packages that instantly make your life easier.
The environment is perfect for data exploration - it saves all the data in your 'environment', allows you to define multiple environments, and your project can be saved at any point, with all the global data intact.
If I want some extra speed, I can create C++ modules from within R Studio, compile and link them, as easily as simply creating a new R script. Fortran is a tiny bit more work, still easy enough however.
Want multicore or to spread tasks over a cluster? R has built in functions that do that for you. As easy as calling mcapply, parApply, or clusterApply. Heck, you can even write your function in another language, then R handles applying that over however many cores you want.
Want to install and manage packages, update them, create them, etc...? All can be done from R Studio's interface.
Knitr can create markdown/HTML/pdf/MS Word files from R markdown, or you can simply compile everything to a 'notebook' style HTML page.
And all this is done incredibly easily, all from a single package (R Studio) which itself is easy to get and install.
Oh yeah, visualisation, nothing really beats R.
And while there are quirks to the language, for non-programmers this isn't really an obstacle, since they aren't already used to any particular paradigm.
As for Python, I'm sure it's great (I've used it a little), but I really don't see how it can compare. R's entire environment is geared towards data analysis and exploration, towards interfacing with the compiled languages most used for HPC, and running tasks over the hardware you will most likely be using.
I think the conclusion of the article is correct. R is more pleasant for mathier type stuff, while Python is the better general-purpose language. If your jobs involves showing people powerpoint presentations of the mathematical analysis you've done,you'd probably want to use R. If, on the other hand, you're prototyping data-driven applications, Python would probably be better.
That said, I really like Julia, but can't justify really diving into it at this point. :\
I would disagree. Python's libraries are really reimplementing R in Python (Mainly Pandas). I find R to be very flexible and especially in the last 5 years with Hadley Wickham's libraries things are concise and very powerful.
Please look at dplyr and see how this new way fo doing R works. Especially with piping with %>%. https://cran.rstudio.com/web/packages/dplyr/vignettes/introd...
Code in R can look like this beautiful code (If you don't code in R and I would expect anyone can see what is happening) This is why I disagree that prototyping in Python would be better.:
flights %>% group_by(year, month, day) %>%
select(arr_delay, dep_delay)
summarise(
arr = mean(arr_delay, na.rm = TRUE),
dep = mean(dep_delay, na.rm = TRUE)) %>%
filter(arr > 30 | dep > 30)
Python has .pipe but I find it strange it goes to the new line before the items.Python Code: >>> (df.pipe(h)
... .pipe(g, arg1=a)
... .pipe((f, 'arg2'), arg1=a, arg3=c)
... )
(df
.groupby(['a', 'b', 'c'], as_index=False)
.agg({'d': sum, 'e': mean, 'f', np.std})
.assign(g=lambda x: x.a / x.c)
.query("g > 0.05")
.merge(df2, on='a'))
There are now methods in pandas to do pretty much anything, so you can chain them together into one easy-to-read manipulation without lots of intermediate variables.Compare scikit learn to other a large number of R libraries with incompatible interfaces. In this respect Python is more regular.
If you need cutting-edge or esoteric statistics, use R. If it exists, there is an R implementation, but the major Python packages really only cover the most popular techniques.
If neither of those apply, it's mostly a matter of taste which one you use, and they interact pretty well with each other anyway.
If most of your job is going to be implementing data analysis techniques that you or someone else has done earlier and putting things into production, then Python will quite possibly be more suitable.
Actually, it is. When someone has only 3 or 4 years to finish their thesis and learning how to program is secondary at best, and they have to do it in a math-heavy department or field, there is no time or use to learn Python.
http://www.oracle.com/technetwork/java/jvmls2013vitek-201352...
Then there is PyPy as well.
I also think they should probably add Julia and Wolfram/Mathematica to these comparisons.
The Purdue project you linked looks quite interesting. Unfortunately, development appears to have stagnated: https://github.com/allr/purdue-fastr
[edit] Another important aspect that Renjin contributes is the packages ecosystem: http://packages.renjin.org/
R's great strength is finding the interesting bits of the data. Testing the Algo. Doing the R&D basically. Better than Python.
Once that's done, why stop at Python? If your game is production, Python will do it, but others will do it so much better, faster, more efficiently.
Unfortunately I can't go into specific details without potentially divulging proprietary information, but broadly most of the issues I've seen in production with R are corner cases involving multithreading with large amounts of allocated RAM (over 100GB), and corner cases involving the data.table package. I've also seen packages that update and break backwards compatibility, although that's less of an issue. The biggest concern we have with R, however, is that the documentation and coding practices for most R packages make small bug fixes difficult without having extensive knowledge of the package code. This is not always true, but it's true enough of the time that we can't afford to maintain much production R code.
The conclusions state what we already know: Python is object oriented; R is functional.
The Last Word appropriately tells us your opinion that Python is stronger in more areas.
The "weekend hack" that was Python, a philosophy carried into 2.x, made it a supremely pragmatic language, which the data scientists love. They want to think algorithms and maths. The language must not get in the way.
3.x is wanting to be serious. It wants to take on Golang. Javascript, Java. It wants to be taken seriously. Enterprise and Web. There is nothing in 3.x for data scientists other than the fig leaf of the @ operator. It's more complicated to do simple stuff in 3.x. It's more robust from a theoretical point of view, maybe, but it also imposes a cognitive overhead for those people whose minds are already FULL of their algo problems and just want to get from a -> b as easily as possible, without CS purity or implementation elegance putting up barriers to pragmatism (I give you Unicode v Ascii, print() v print, xrange v range, 01 v 1 (the first is an error in 3.x. Why exactly?), focus on concurrency not raw parallelism, the list goes on).
R wants to get things done, and is vectors first. Vectors are what big data typically is all about (if not matrices and tensors). It's an order of magnitude higher dimensionality in the default, canonical data structure. Applies and indexing in R, vector-wise, feels natural. Numpy makes a good effort, but must still operate in a scalar/OO world of its host language, and inconsistencies inevitably creep in, even in Pandas.
As a final point, I'll suggest that R is much closer to the vectorised future, and that even if it is tragically slow, it will train your mind in the first steps towards "thinking parallel".
I've grown to appreciate R, especially its plotting ability (ggplot).
But a few weeks back he asked me how to do some kind of data sorting / manipulation in R. My answer was that it was a 10 line Python script and I gave him the code. Alas, he couldn't figure out how to save the script and run it from a command-line.
You can't underestimate at how important Rstudio is to the popularity of R for non-programmers.
This. Most programming IDEs show the code but hide the data. Excel shows the data but hides the code. RStudio is awesome because it shows both the code and the data.
That being said - all the serious math/data people I know love both R and Python...R for the heavy math, Python for the simplicity, glue, and organization.
He says "In R, there are packages to make sampling simpler, but aren’t much more concise than using the built-in sample function" but using caret is more concise.
Added: Later in the section on random forests he says "With R, there are many smaller packages containing individual algorithms, often with inconsistent ways to access them." Which is why you want to use the caret package as it makes accessing many machine learning packages consistent and easy.
Python: There should be one, and only one, preferable way to do things. Though this may not be obvious at first.
R: Every author has a different style of doing things, reflecting in the code.
As for the comparison in general: You can call R from within Python. So Python is at least as powerful as R. The rest (BeautifulSoup, Compression, Game development etc.) is icing on the cake.
What features or workflow does R or Pandas/Numpy offer to manufacturing that Minatab & JMP can't?
I don't know anything about Minitab/JMP scripting myself, but my understanding is that R is generally the most intuitive of all the aforementioned (although that would basically boil down to individual preference).
Here's a review including Minitab and R that might be of interest: http://www.prostatservices.com/statistical-consulting/articl...
The equivalent comparison should be R+dplyr to Python+pandas.
Base R is quite verbose and convoluted compared to using dplyr. Likewise data analysis in Python is painful compared to using pandas.
An alternate (simpler) implementation of the rvest web scraping example is at https://gist.github.com/jimhester/01087e190618cc91a213
It would be even simpler but basketball-reference designs it's tables for humans rather than for easy scraping.
End of the github for rvest:
Inspirations
Python: Robobrowser, beautiful soup.IMO, R's system is actually more powerful and intuitive -- e.g. it is fairly straightforward to write a generic function dosomething(x,y) that would dispatch specific code depending on classes of both x and y.
R's value is in the implementation of its libraries but there is no technical reason a really OCD person couldn't implement such high quality of libraries in Python.
> data(iris)
> library(caret)
> data(iris)
> idx <- caret::createDataPartition(iris$Species, p = 0.7, list = F)
> summary(iris$Species)
setosa versicolor virginica
50 50 50
> summary(iris[idx,]$Species)
setosa versicolor virginica
35 35 35Most R tasks that people use exit. Typical data science task is: gather data, apply an operation over said data, analyse results.
python < world > csv
R < csv > analysisto me, R was a waste of time and I really dont understand why its so popular in academia. if you already have some programming knowledge, go with Python + Scipy instead
EDIT: R is even more useless without r studio, http://www.rstudio.com/. and NO, dont go build a website in R!
That may not be what you meant, so I haven't downvoted yet, but it doesn't seem to be an attitude that is helpful for the conversation.
What I meant to say was that I helped my wife during her master thesis (~6 months) with R, in addition to spending an hour in one of the classes.
Her teachers also were novices of both R and Excel, and we had several issues with everything from how R processes csv:s, to just figuring out the proper syntax to have R do what we wanted.
Sorry if my comment wasnt helpful, i was merely attempting to add some reflections from personal experience to the discussion.
As a side benefit, the first time I tried dabbling in Julia, I was pleasantly surprised to have a familiar mature environment work with it out of the box.