A book to learn R and Python in parallel for Data Science
github.com
github.com
While previously Shiny was primarily deployed through RStudio's solutions, there are now open source initiatives such as ShinyProxy, introducing Kubernetes as an option for deploying Shiny applications. The latest iterations of Shiny related libraries are facilitating automated testing and deployment. These developments allow companies to use Shiny in production, but it has to be said that the R ecosystem is not as developed as Python's from a traditional software development perspective.
By Dash for Python, you mean the one from Plotly? https://plot.ly/products/dash/
Thank you for sharing ShinyProxy!
ShinyProxy is amazing. It is pretty easy to setup, but does require quite some specialized knowledge compared to the RStudio solutions.
Awesome!
https://dash-app-dx9g2r0la6-8000.cloud.kyso.io
It was also really easy to make it, maybe 250 lines of python in total
(guide to making this app is here: https://kyso.io/KyleOS/creating-an-interactive-application-u...)
R has such a huge library of software now, that it has gone far and wide outside of statistics and analytics - any kind of application can be built in R now.
In fact I see no point in using Python for mathematics or number crunching any more as R has it all and performance critical parts can be rewritten in modern Fortran for very high speed and made available inside of R transparently.
As you gain experience building applications, at some point you'll learn that you were wrong. Or you won't, in which case I feel sorry for whoever inherits your reinvented-wheel codebase.
The various software I've written is small, fast, light, with minimal dependencies and it's easy to install because I deliver it as OS packages for the operating systems I support; my users tell me they are happy. The memory requirements are miniscule and the software lightning fast. Its size is measured in kilobytes, not megabytes or gigabytes, which means I must be doing something right. The manual pages often exceed the software in size and are brimming with examples. I pay extra attention to being backwards compatible when I implement changes and enhancements. Regressions are non-existent.
So I know I'm right and that using a bloated "framework" would have been one of the stupidest things I could have ever done.
As for re-inventing wheels, I use what comes with the OS and leverage what's already there; I've purposely not implemented any algorithm re-implementations of my own, although I easily could have. I'm neither dumb nor stupid to go re-inventing wheels, in fact that's one of the reasons why I hate webshits' frameworks. They don't call them webshit for no reason.
This isn't true at all...
Also all advance statistical books are either SAS or R. If it's R then there is always a package that the author created.
Just look at Chapman & Hall/CRC or Springer publisher and look at their books.
Go here: https://www.jstatsoft.org/index
Count the number of R packages in those papers versus Python.
I don't even need a source. I'm a statistician and I'm going to get a paper there and publish a R package for my master thesis.
Although I haven't had as much reason to use base R. For more ML-related tasks I do go back to Python.
The overall api of tidyverse packages is such a joy, and recent improvements in purrr/tidyr allow me to construct nested data analysis workflows I couldn’t even dream of in python.
https://forcats.tidyverse.org/reference/fct_lump.html
There's also the data.table package for this kind of data work, which is maybe less used but seems to have better performance.
For example, I know the recommended pipe in R is magrittr's %>%. I have no idea what the respectable pipe library in Python is, or even if there is one.
I wouldn't even know where to start finding all the tidyverse equivalents in Python. It isn't as organised and obvious as the R statistics community.
On the other hand Base R is the worst. Disgusting language.
ggplot2 : plotnine is quite good ggplot2 clone based on matplotlib. I feel like ggplot2 is a bit better and more complete, but if you want to do something that isn't supported it's harder for me to hack than matplotlib.
Tidyverse: To me, ggplot2 is the only essential part of the tidyverse. Lubridate is also good. Most others seem like semantics and syntax sugar. I prefer data.table, which is similar to Pandas. DT is super fast but imho Pandas has a more intuitive and consistent API (and if you want a speed up for large N then dask might work).
I use both R and Python on a regular basis. I choose Python for lower-level stuff, automation, parallelism / concurrency, and R for bespoke statistics. I use both for everyday statistics and plotting, but I feel that R has light advantages. I feel like if you're comfortable switching languages there are good reasons to use both. It's also important for me because I work with different teams that have different practices and preferences.
There's a bunch of comments below which can be summed up with 'use R because <package name> doesn't have a direct python equivalent' but they're all missing the point that the Python data science ecosystem is evolving at a much faster pace than R and will completely supersede it in a few years.
R, like SAS, is a tool for non-programmers. And there it shall remain. The only demographic where R makes sense long term are pure mathematicians/statisticians who are not proficient in programming. But that demographic is rapidly declining in size.
The point is R is a very good language for statistic because of the packages not data science. Data science can do their own thing it's okay. It's also okay for data science to use statistic models from statistic too.
> R, like SAS, is a tool for non-programmers.
I respect and love data science and machine learning but this behavior of generalization is terrible. There are many wondeful programmers contribute to R and uses R as I am sure there are many wonderful statisticians that use Python. They're just tools.
> And there it shall remain. The only demographic where R makes sense long term are pure mathematicians/statisticians who are not proficient in programming. But that demographic is rapidly declining in size.
What is up with these generalizations? R is not going anywhere in the statistic community. It's doing fine. Also from my experiences in academia most math people use matlab and if any R.
It's okay to have both R and Python doing their thing.
There is no need to conflate data science and statistic or have this weird tribalism.
R has nothing going for it except a rapidly dwindling number of packages that don't yet have a direct python equivalent. It doesn't make sense to invest time into R if one already knows python unless one specifically focusing on academia pure stats type stuff.
Even then, the incoming generation of undergrads are increasingly proficient with programming and are shying away from R the same way that they shied away from Matlab after scipy matched it for 95% of their tasks.
Ha! I knew it! So it is familiarity with Python then!
R does have something else going for it: phenomenal documentation and consistency. Replicating R's thousands of available libraries will be a gargantuan effort. It is cheaper and more efficient to master R.
This is not a true statement.
Here are the data that goes against this statement.
1. https://www.r-bloggers.com/on-the-growth-of-cran-packages/ 2. https://blog.revolutionanalytics.com/2017/01/cran-10000.html 3. https://www.r-bloggers.com/rs-remarkable-growth/
From 2015 to 2016: ~6,200 to More than 8,000 in April, 2016
From 2016 to 2017: CRAN now has 10,000 R packages.
> Even then, the incoming generation of undergrads are increasingly proficient with programming and are shying away from R the same way that they shied away from Matlab after scipy matched it for 95% of their tasks.
This is a generalization.
So far you've made opinionated negative generalization with no data.
Python is great because it learn from Matlab and took many great ideas and inspirations from Matlab. But I'm not going to make sweeping negative statements about Matlab or pretend to know how it going when I don't have enough data or experiences in it.
What? The R ecosystem doesn't provide meaningful out of core capabilities, nevermind the ability to handle anything approaching 'massive amounts of data'.
-- Would sure love to know why an agenda-less factual comment is getting downvoted.
Column stores are standard in any analytics pipeline today. They make up Python's Pandas, R's dplyr, and Java's DataFrame. How or why does R stand out for 'massive amounts of data'?
R does not have have meaningful out of core compute offerings that compare with something like Dask.
R does not at all have cluster compute offerings that compare to Dask Distributed.
If you want to know what real performance looks like, check out Python's cudf which will shortly fully match the Pandas api. That raytracing example you linked would run at interactive rates with cudf, I really don't see any basis for perf arguments in R's favour, and 'massive data' arguments are laughable here.
Whatever advantages R has, perf or scalability are definitely not amongst them.
I don't see how the "GPU DataFrames" provided in cuDF would enhance a raytracer in any way.
Bonus: modern Fortran is a joy to develop in, far more fun than Python. And you get to compile to machine code, either for a processor or a GPU.
I still occasionally use pandas with seaborn when it's not worth it to switch out to R. I don't think it can match the tidyverse+ggplot combo for quickly exploring and making beautiful plots. But this discussion has inspired me to do some googling and it seems like some people are using tidyverse-like workflows in pandas (https://stmorse.github.io/journal/tidyverse-style-pandas.htm...). Doesn't seem quite as smooth but I'll definitely be trying it out next time I'm working in pandas.
Some of the dplyr elegance comes from the flexible evaluation mechanism in R, whereby mutate(data, col1+col2) works because the second arg is evaluated in an enriched environment. Python eschews this kind of macro-like extensions because, my guess, tampering with evaluation makes a lot of other things complicated (for instance, forget replacing args with their value, that doesn't work anymore). I think the author of dplyr himself in later work has promoted the use of the ~ operator to explicitly block eval of an argument and at least make these departures from regular eval explicit. That means dplyr is ahead for interactive use, but for programming you have to switch to a separate API (the underscore "verbs") and that makes the transition from interactive work to coding a bit steeper. It's all trade-offs, and I am not saying that I know better than either the pandas or dplyr authors.
As to ggplot, if you believe the future of statistical graphics is in-browser and interactive, you should take a look at altair for python (I myself created a small extension to it called altair_recipes). It's based on vega, like ggplot anointed (but not quite ready) successor ggvis and uses the grammar of graphics (or on interpretation thereof) like ggplot, with extensions to interaction. Simpler than D3 by most accounts.
Although I like R and often use R to quickly order tabulated data, there are a few things to take into account that in recent times are building a strong case for me not to use R habitually.
Development in R is frustrating. If you don't need to do dev, then on this point you are home free. Testing things that you deploy in R is not simple.
Scripting in R can be frustrating. I have a script that traverses Excel files and using tryCatch() is just so much more complicated with it being a function. In Python the try-catch functionality is part of the design syntax.
There are scenarios where R is better. If you are in actuarial science, research or academics then often you'll find R libraries that just work.
R treats tabular data with grace. Everything in R is an array.
The takeaway for me is that I should use R less and Python more. I personally can't deal with something like tryCatch() being overcomplicated, but for people who don't do dev anyway and maybe need to analyse DNA sequences for a living, R can be rewarding. For me: the ggplot2 library is great; stay away from Shiny and dev in R.
It tries to do functional programming, but the documentation is not satisfying. The responses and behaviour is perplexing.
I spent around 5—10 hours trying to get a Shiny GUI to work and eventually got to the conclusion that 1) if you want a big project do all the frontend stuff in something else, like JS and 2) if you want a small project try something established (I am not advocating, it's just an example) like Power BI.
Rather than R vs Python I hope one of two things happen. Either both languages get replaced by a 'better' ML language eg Swift / Julia giving us users a 'turtles all the way down' experience and removing the reliance on complied packages. Or, second option, they get relegated even further into being nothing but glue between some common data formats specific to the type of work found in DS allowing you the user basically a choice between syntactic-sugar of one glue-language versus the other. Something like Apache Arrow springs to mind but I'm not sure where they are at the moment
If anyone is interested, I also made a 'Learn R by Example' project which attempts to teach R through code comments: https://github.com/photonlines/Learn-R-by-Example
There's a good explanation here: https://stackoverflow.com/questions/1741820/what-are-the-dif...
e.g.
divide = function(x, y) {
return(x/y)
}
divide(y = 2, x = 1)
divide(y <- 1, x <- 2)
These two calls give the same result, as the second results in assignment and then passing the argument by position. Other than this case, they are exactly interchangeable.Teaching and executing are two separate skills. Fun little anecdote, in high school I had this AWFUL science teacher. He would literally just have us watch Crash Course videos to get the concepts. Turns out he was a relatively distinguished scientist himself..