How the R-project is taking over statistical analysis software
sites.google.com
sites.google.com
Open-source is great at hill-climbing, where there are clear directions for improvement and especially for features that are obviously needed by users (provided the structure of the project is sufficiently modular to facilitate it), by tapping the collective intelligence of users.
It's not great at "hill-hopping": originating radically different products.
Open source only seems to win in domains in which it makes sense for companies to share work in order to compete at a higher tier of functionality.
The only way it realistically can realistically happen if is the commercial product is not being improved (i.e. differentiated) any more. Your example is one of these - standardization is related to commodification.
I suspect also that user-facing apps are easier to keep improving, because the user is right there, and always has more needs that could be served (e.g. text editors will evolve til they can read mail; those that can't will be replaced by those that can). Non-user facing apps tend to be defined by their environment, rather than by users - although, any component that creates a benefit that the user wants more of will keep being improved (from Clayton Christensen). e.g. databases, CPUs.
I don't think it makes sense to make generalizations about where "open source seems to win". Things are changing too fast; the circumstances that made it possible for Mozilla to beat IE in the mid-2000s no longer exist, for example.
But I agree with you and think that in the long run, open source is a big winner in infrastructure software, and generally just that.
R's integration with free data sets is crazy. For quick analysis it is great. For development Numpy/SciPy/Matplotlib/PyCUDA is more suitable.
One other potential factor: a lot of this software is driven by academic use, either because academics used it or that's where people were first exposed, and academics often receive large discounts.
Today's Octave installs do not need more than click-click-ok-done or apt-get install octave.
> for which the open source clone Octave is a bad joke
I think you're not giving Octave enough credit, considering that they have only a few part-time developers, and that nobody does sponsor them, they have accomplished a respectable amount of functionality over the last 20 years, it is extermely unfair calling them a "bad joke".
Of course with those limited resources they are not able to match the output of Mathworks, but what they can do as of yet is usually more than universities teach, and _still_ many departments are kind of married to Matlab, only mention Matlab to students, give only matlab examples, matlab labs, matlab exercises, etc. Also very respectable people like Gilbert Strang, who gave the MIT basic Linear Algebra and Computational Science and Engineering classes, seem to have enough vested interests in Mathworks to not even mention Octave to students briefly as something they can download and work at home. Octave is extremely powerful and capable for what you pay for it, and deserves at least a mention.
It is probably not different at other universities and other departments. Several professors I had to deal with were similar, either not even aware that open source packages like Octave, Scilab, Maxima, Scipy exist at all, or extremely faithfully married to companies behind proprietary packages like Matlab/Maple/Mathematica.
http://www.cs.utexas.edu/~EWD/transcriptions/EWD12xx/EWD1283...
The grammar is pretty spot-on although there is usually some release latency when Mathworks changes it (obviously since their plans are not made know. Ahead of time).
Octave even has Matlab source level compatibility for mex files although they are slower than octaves own c interface.
If you start drifting away from Matlab core needs into the specialized add ins Mathworks provides (simulink, financial packages, etc) then Octave can't help. If you need those then I find that Matlab is rarely the tool for the job either (you just don't know it yet ;))
R is sick though and I am alwalys pleasently surprised at what clever people are doing with it. Octave on the other hand shouldn't be used until someone writes a proper interface and decent graphing. MATLAB is a million light years ahead of Octave in that regard
https://sites.google.com/site/r4statistics/_/rsrc/1318535062...
My only complaint is the awful default IDE, which can be mitigated to a large extent by scripting elsewhere and source()ing the script, and some odd edge behaviors including the mystifying row names of dataframes, the difficulty of dropping unused factor levels from aggregated or sliced data (another dataframe issue), and the perhaps unnecessary obscurity of some of the plotting functions (although holding R responsible for the lattice library is unfair).
All that said, for a free tool, it's extraordinary, and the authors of the base language and the many packages that I use have my gratitude.
http://www.frontiersin.org/human_neuroscience/10.3389/fnhum.... http://www.mitpressjournals.org/doi/abs/10.1162/jocn_a_00089
I'll need to give Stata a look, too.
> Robert A. Muenchen is the author of R for SAS and SPSS Users and, with Joseph M. Hilbe, R for Stata Users. He is also the creator of r4stats.com, a popular web site devoted to helping people learn R. Bob is a consulting statistician with 30 years of experience
Disclaimer: I hate R's syntax, but my company's analytics group uses R for just about everything.
I've complained before that Octave is the wrong solution to the Matlab problem, and if you aren't attached to one of the many fine Matlab toolkits, you're likely better served translating to a more expressive language, like Python+Numpy+Scipy.
The biggest difference between Matlab and Octave is JIT compiler in Matlab, which does incredibly good job at vectorizing simple (or sometimes even not-so simple) loops.
I think it's fair to say that Octave performance is very close to a Matlab in a pre-JIT time.
There's also a huge difference in toolboxes, profiling, sparse matrix operations, parallel computing and many-many more. In these areas I'm afraid Octave is light-years behind Matlab.
However, you still can do a lot of useful simple stuff with Octave and it's free! Matlab-like syntax is really, really cool then it comes to vectorized operations. So probably these two reasons determined Andrew Ng's choice of Octave as a main environment for ml-class. Huge win for Octave I guess. This might spur some interest in the development, attract new people to the product. I think it's a well-deserved success for John W Eaton and other people who develop(ed) Octave all these years.
As you note, the Matlab profiler is very nice. You can zero in on the 80% of the 80/20 tradeoff very fast, during your usual development cycle. It's as simple as:
>> profile on >> do_something >> profile report
and you get a nice graphical/textual report on time usage in everything do_something called.
This is not true. They strive for Matlab language compatibility, but none of them refers to Octave as a "Matlab clone", nor are they working on cloning Matlab, nor was the project started to become a matlab clone. It is like calling Linux a "Unix clone".
http://radar.oreilly.com/2011/10/oracles-big-data-appliance....
It's probably more an issue of easily pre-filtering/aggregating the data before analysing it with R. I like this approach of moving the calculation to the data, but we must be very late on the adoption curve if Oracle are doing it already.
There are two ways to deal with that, one is to load datasets through SQL database (using a SQL library) which IMHO is a "dirty hack". The other (what I usually do) is to load the huge datasets in STATA (or any other stats package) and filter the data to get a set that is small enough to work with R.
Other than that, the available libraries in R are crazy good. for example stuff like Approximate Bayesian Computation or survey analysis (considering weight factors) is straightforward with available libraries.
There are a huge amount of available libraries (thousands!) of variable quality thanks to the open nature of the project. But commercial software has problems too, especially with new and niche products. And when something goes wrong in those cases, you can't see why for yourself. Worse, other independent experts would not have the chance to either.
My company shuns R (although I personally like it), primarily because of this issue. If we need to run a rare or uncommon statistical procedure, it is a lot easier to trust the SAS procedure, rather than an open source R package written by some grad student.