How R Took the World of Statistics by Storm
statisticsviews.com
statisticsviews.com
I'm taking a break from Octave, but I plan to come back to it and take Matlab head-on, not chip away at the edges.
What are they?
Anyone that isn't using those could probably switch to octave with few issues but last time I checked some functions were starting to differ (see eig)
Octave is nice when you really need to have matlab compatible files but don't have a matlab license but now everyone that I know in the sciences if they are not using matlab and don't need to use fortran/c/c++ they are using python because why keep the matlab syntax if you can gain so much by switching to another language.
EDIT: Also, although it's not opensource, I had used Mentor SystemVision eons ago in college. It now has a free web interface. I haven't tried it, but honestly, what could be worse than DxDesigner ;-)
Matlab is a crappy language, but a productive environment.
Also, mathematicians and statisticians think functionally and the general attitude in python is to do object oriented programming while R is strictly functional programming with a little bit of object programming.
However, in my experience, for the data munging required as a preliminary to the analyses, R is worse than bad. It's as if satan himself designed a language.
I find that what then happens is this: data scientists/statisticians/[your favorite word here] become reliant on programmers to clean/format the data to do the analyses.
This is all fine, but those same scientists are then put off learning python, where they could do all of their own munging, and probably 95% of the analysis they need to do, and where they could further add value by writing programs that are easier to production-alize.
Job security for those who know how to write production code, I guess.
I don't know anyone who considers themselves a "data scientist" of any sort that doesn't view their job as 80% or more data wrangling/munging/cleaning.
I write production ETL processes in R at my current job. AMA.
I'm always interested in learning a new tool, though.
I have largely avoided ts, zoo, etc where possible. Time series stuff seems to have a lot of specialized tooling all of which tends to be much more strict about data structure than I'm comfortable with for my flow.
[1] http://opiateforthemass.es/articles/james-bond-film-ratings/
BUT perl + R was a really nice combination for a while.
A lot of package are built from matlab and not for Python, if you are a PhD student you want to use those package, not write your own, if you really really need to write some software you want to write it on top of something that already exists...
It is really sad, but it is the reality...
In grad school, if I wanted to do my own EEG connectivity analyses, I could just include the Signal Processing Toolbox, the Stats Toolbox, and crunch my own numbers. Or, if I wanted to do a more standard analysis of my fMRI or EEG data, I would turn to the world's most popular open-source toolkits (SPM and Fieldtrip), both of which require... you guessed it, Matlab.
The only place I ever found Matlab's libraries deficient for my needs was in machine learning. (I ended up doing an SVM-based spotlight fMRI analysis in Python.)
There's a lot of lock-in and quality toolboxes around Matlab, Python/Julia won't knock it over yet, though I wish them the best of luck.
R was much more programmable that the other systems (except S) in that time -- while the R language is not pretty, the scripting languages in SAS, Stata etc. were much worse. So R provided an improvement in people's workflow. Whereas Octave is just a free version of the same language as Matlab (with some small improvements).
So, switching from SAS, Stata, SPSS etc. to R provided an improvement in productivity, but switching from Matlab to Octave does not.
This is not the reason for why Octave has not overtaken Matlab. Maybe I should write a blog post about it.
The bigger issue is that while R is liked by statisticians it lacks many of the features for the software development. We run across difficulties with logging, version control of packages, speed, size of docker image, build time etc. But, with these drawbacks I keep coming back because I develop faster and better in R.
I recently had to loop through 1.3Gb of data (5000 files) and merge just one column from each file into a new dataset. It did so in ~2 hours. Yet the loop was just ~5 lines of code.
I wonder if you tried doing things like:
* preallocate a list, then do.call(cbind, your_data) * Same as above, but with some of the faster alternatives to cbind like dplyr::bind_cols or data.table::cbind * Use data.table, which has far faster joins than base R (so does dplyr) if you were doing a true merge/join
If it was truly just adding a column rom each file together into a file, these kinds of tasks are much better using UNIX tools, in my experience.
Another example is the immutable structure that causes R to be a memory hog. Creating copies of data everywhere. But, again if you plan well and execute the 'best' solutions you can avoid the giant pitfalls but will rarely ever beat a equally well written python equivalent.
packrat has helped a lot with version control of packages, but it still doesn't quite feel like the right solution.
I've been really impressed in the last 5 years how far R has come in these areas though, so like you I keep coming back. By the time I start getting over the learning curve other places, R seems to have developed better tooling for what I want to accomplish anyway and I can come back and right cleaner, clearer, better software faster in R.
I have some programming background and really would like to get into statistics. Should I do some R tutorial and throw my weblogs at it to see what I can do? Or is there some awesome learning resource you could share?
http://www.meetup.com/Dallas-R-Users-Group/pages/R_Helpful_L...
Octave replaced Mathlab.
Python based libraries are somewhere in between.
Julia with Jupyter will probably replace Mathematica, LabVIEW and Mathcad (and unify all of the above) with a powerful native language and environment.
What does it offer in terms of controls systems/realtime, GUI building, and data flow?
R replaced S and S-PLUS, not SPSS. SPSS is still around as a light, user-friendly stats tool.
And Octave is nowhere near replacing Matlab, not by a long shot. It's the complete opposite story as R/S-PLUS.
Source: was in grad school for cognitive neuroscience. Saw Matlab everywhere. Saw SPSS here and there. Saw Octave nowhere. Briefly looked at Octave and stopped as soon as I realized all of the packages everyone used required Mathworks toolboxes.