SciRuby
sciruby.com
sciruby.com
I'm in the middle of writing a Stats library for Ruby[0]. Maybe we can join forces?
Our distribution gem -- well, technically, Claudio's distribution gem, but used by SciRuby -- has some Ruby code in it derived from C code in the GNU Scientific Library. Being a GNU library, GSL is licensed under the GPL.
Interestingly, the GSL-derived code is only utilized if the user does not have libgsl installed. And my understanding is that code which uses libgsl is not technically a derivative work, and therefore not required to be GPL'd itself.
I suppose one possibility is to abstract the GPL'd code into yet another gem (distribution-gsl?), which is itself licensed under the GPL.
Non-software shops who may be interested in using a piece of software may balk at using GPL3'd software because they don't know what the legal ramifications are of failing to comply, or the knowledge/wherewithal/processes to do release the software that they're using.
Talking to people about BSD/MIT is really easy: "You can do whatever you want with it as long as you retain the license and copyright".
At the risk of getting into a FOSS license debate, i'd like to think that FOSS contributors do it out of a motivation other than license restriction (and hell a lot of people still rip libs off, even when they are GPL'd!).
It is also the case that this is basically a library, and many who would have no problems using/contributing on a GPL application will balk when it comes to a library or framework.
In other words, what could we do to provide another arm-twisting mechanism to force publishing authors to release their source code? We wanted to facilitate openness among people who might join the Ruby community by way of SciRuby, as opposed to those who might join the SciRuby community by way of Ruby.
Whether or not we can actually enforce release of source code is an open question. Certainly many journals and funding agencies have enormous problems here.
What about a joint license? For example, is the following practical? "If you publish your work in an academic context, SciRuby is GPLv3 for you. If you do not publish your work in an academic context, SciRuby is MIT."
The two technical justifications (objects all the way, and enumeration) in the FLOSS interview [0] are both arguable and not nearly convincing enough to justify further fragmentation of the open-source science ecosystem. If I want cleaner semantics, s/python/ruby is at best moving sideways - for the sake of a few keystrokes? Ruby is slower both in the interpreter itself and in the lack of f2py,Cython,Numexpr,PyCUDA,weave (even Theano sometimes).
If I need a real change, I'll use Ocaml or Clojure, and gain speed from the change.
It seems like a waste to discard (or attempt to replicate) the 10s-100s of person years represented by SciPy and the ecosystem including f2py, Cython, MayaVi, IPython (not just a REPL), Pandas, Chaco, PyCUDA/OpenCL, and SAGE - to name a few.
Are there any other, better reasons to want to build an ecosystem from scratch?
Some arguments: * Ruby is expressive and flexible in ways that python is not. For example, Rubyvis is flexible enough (or similar enough to javascript at least) to essentially accept protovis (javascript) code directly. I don't think python can do this. * Ruby uses blocks/enumerators instead of 'for' loops. How much programming involves enumeration of one kind or another? * len(array) vs. array.length
Coding is not just getting the computer to do what you want, it is also how you think about it and the form that it takes. 'len(array)' vs. 'array.length' may not matter to most, but it matters to some of us.
The great news is that we can still use scipy when we need to. I'm betting there is room for both projects, especially considering how small sciruby is at the moment.
Ocaml and clojure are great, but have a steeper learning curve; getting non-programming scientists to contribute is far easier in a language like ruby or python.
A beginning-programmer scientist who wants to start writing code to solve problems currently has to 1) code in python or 2) learn ruby and python/scipy or matlab or R in order to do some scientific computing. SciRuby means (eventually) that for most things, novices only have to learn ruby. It is hard to overstate the importance to new programmers of being able to use just one language (at least to start with).
If you like python over ruby, this is easy. If you like ruby over python, it gets old piping all your data over to a python script to use the basic features of scipy.
For the record, NumPy predates PDL by a smidge, and the existence of PDL is more of an argument against SciRuby, considering the cultural similarity and continued strength of BioPerl. As for C/C++/Fortran wrapping - there's a bit more to it than syntax efficiency, plus wrapping is often semi-automatic and leverages multi-language, science-ambivalent toolkits (ie SWIG or SIP). However, one nice consequence is the fact that via buffer wrapping, the NumPy array has become an efficient common currency for a huge number of legacy libraries.
IMHO, the advantages you have cited pale in comparison with the task of reimplementing 15 years worth of work for what is essentially unity gain in code style (+/- 2% depending on your flavor preference). Regarding that code style, you may be missing the forest for the trees: the inflexibility of Python is a small price to pay for community cohesiveness, and the resulting multiplier effect is non-trivial. Put another way, the time I've spent attempting to read Perl code-golf leaves me very leery of Ruby.
My basic argument though is not that Python is superior or coexistence impossible, it's that every person-year spent reinventing a mature system that is far beyond 'good enough' is a person-year that could be spent advancing the state of the art in scientific computing by building on Theano or improving SAGE - or preferably, doing real science.
You seem to be saying that ruby is python, just with perl's inconsistency and unreadableness. The syntax differences between python and ruby (more than 2% IMHO) amount to very large differences in code organization. The ruby community places a high premium on brevity and clarity, and ruby's flexible syntax facilitates this. Part of the reason monolithic code bases are more rare in ruby is because we tend to do more with less code. We are talking about two very different forests.
We weren't going to be working on Theano or SAGE anyway, just doing basic science computation to solve real problems in our fields. That's really the problem, we find ourselves quite productive doing science in ruby at the moment, despite its relative immaturity in this area. It's hard not to imagine what we could do with a few more foundational tools. It really is a small cost to enable us to be able to do science in ruby. Plus, building scientific computing libraries is good fun. Every community should have the chance. As Abe Lincoln suggested, 'Let not him who is houseless pull down the house of another; but let him labor diligently and build one for himself.'
Is it so terrible to want to do science in the language you most enjoy and are most comfortable with? To continue with the maladroit quoting of Abe Lincoln, "it's best to not swap horses when crossing streams."
I made the point further down, but will restate it: if ruby hadn't persisted in existing (the nerve!), would django exist today? I find it hard to believe that this is a zero-sum game, and I wish python and scipy every success. Perhaps the most valuable scipy contribution we could make will come by making sciruby something worth borrowing ideas from?
SciPy is great, but it's clearly best for programmers that have a slight scientific bent and can't stomach learning the existing scientific tools (which are admittedly a bit difficult to combine with modern software engineering). There are some great ideas in SciPy, but a broader set of influences is essential to making a great scientific toolkit.
I'd love to hear more details about the deficiencies, or how it might be more influenced by those.
I agree that the ideas of statistical processing in R are absent from Numpy, but Pandas is attempting to remedy that.
What do you find missing in SciPy?
The sum of anecdotes is not data, but
It might be the opposite: people who know the pain to work with this tools move to Python for complex projects if they can.
Great point about sugar water. Totally on point.
I want to interoperate with C code without disgust. I actually want to write my computation kernel in C, inline with my general high-level code. Numpy does that and I LOVE it for that.
I want to be able to use libraries that are not provided by Mathworks. Serious GUI programming using PyQt, for example. Or, say, a decent XML parser. Or maybe some JASON importing.
I want to be able to run my scientific programs without having to wait for a huge Matlab installation to start up and I want to be able to use my terminal properly.
I want to not have to run X11 on OSX for christ sake. (Though apparently, this has been alleviated to some extent in the latest version. Anyone know first hand?)
Oh, and I don't want to fork over 3k bucks just to write some simple signal processing stuff. (And I don't want to pay upgrade fees every year.)
But then, I don't have an infinite budget, I am mostly interested in signal processing for audio signals and I certainly have more of a programming background than a science background (though my formal education would have me believe otherwise). Also, I am not much interested in Simulink and my latest version of Matlab is of 2007 vintage.
That's why I'm trying to use Ruby where possible, even at the cost of a small productivity hit. The benefits of others being able to read my code far outweigh the few extra minutes it takes for me to do something (and in many cases, the sheer brevity of Ruby as a language means it's faster, simply because it's less typing).
I'd love to see SciRuby become a more useful project, and I'd love to contribute. Unfortunately, they don't make it especially easy to get involved -- the mailing list points people to the roadmap, but it's not at the level of detail where someone could jump in (and the component gems don't seem much better), so it's a bit hard to know where help would actually be useful.
The reason is, they like ruby so much that they want to use it for number crunching too, and not to have to use a less appealing language. They want to build an ecosystem so other people can join and contribute and grow together and maybe outgrow python. They want ruby to win so much that they are willing to work on duplicating a framework existing somwhere else.
1. Because of its consistent object oriented design and because everything returns a value, chaining is natural in Ruby.
2. Avoid index errors and for loops with powerful block Enumerators.
3. Scientific data and services are moving to the web, and Ruby is a great web language (although it is incorrect to call it 'just' a web language).
4. The Ruby community is highly innovative and dynamic, so we can collectively generate solutions quickly.
(see the interview for more explanation) Some of these points are more vision than reality at this point, but we think ruby has great potential as a science language. The other reason this makes sense is that folks are already doing science in ruby (and have been for some time) and many tools already exist--we are building on a good foundation and just hoping to extend it to make things better/simpler.
2. Python has list comprehension, iterators, and generators, which also allow you to loop without for loops or explicit indexing
3. Python also has lots of web frameworks/libraries.
4. Are you saying the Python community is not (as) innovative?
I'm not saying this project is completely without merit. I know that many people prefer Ruby's syntax and preferences to those of Python. But I think it is disingenuous (and somewhat insulting to Python developers) to claim that this project has any potential benefit beyond allowing Rubyists to do scientific computation in their preferred language.
However, due to the special nature of python's builtin datatypes (str, list, int, float, tuple, dict, etc.), you will not be able to use chained_class on those classes.
This helps to make my point about ruby's object model. Of course, ruby takes a speed hit for it, but it is more consistent in this regard.1. (reply) I know it is possible, however, in practice it is easier and done more often in ruby. The idea permeates ruby thinking much more than in python. Even 'if' statements return a value, and rubyists often use that.
2. (reply) yeah, list comprehensions are cool, wish ruby had 'em. The original point was more directed to R and matlab code. Still, the way in which enumeration is done in ruby looks and feels quite different than python, even if the same effect is achieved. [1,2,3].each_cons(2).map {|x| x + 3}
3. (reply) True, but if ruby hadn't persisted alongside python, would django exist? sciruby may do things differently than scipy, and that just might make python/scipy better down the road.
4. (reply) To be clear, I have the utmost respect for the python community. How about saying that they innovate in different ways? Python's emphasis on having one preferable way of doing things tends to yield well engineered systems, and innovation occurs easily in layers as a result. Because of ruby's flexibility, Rubyists are more prone to rebuild core functionality in different ways. Depending on how much you care about syntax, that's either a colossal waste of time or quite useful.
Finally, if you don't care about the syntax of scientific computation, then yes, this only allows rubyists to do computation in ruby. But, despite their many similarities, ruby and python are fundamentally different in several ways. To suggest that we might be able to do scientific computing in somewhat new and interesting ways is not meant to be insulting to python developers, and we make the claim with the deepest sincerity, and humility. Like python, ruby possesses unique strengths, and we hope to bring these to bear on scientific computing as best we are able.
Why? (just curious) SciPy is mature, popular, and has already been heavily peer-reviewed. Then there are R an Matlab and ...
What does SciRuby bring to the table that makes it stand out from the rest?
1. better chaining of commands 2. blocks and enumerators 3. integration with rails and other web services 4. a dynamic community
The screenshots of protovis/d3 look very promising, I'll have a look at it. The last time I needed a JS charting library I went with Highcharts, as it had somewhat better support for the run-of-the-mill chart types I was using in my project.
Protovis/d3 take a different approach, also focused on a similar Grammar of Graphics like ggplot but primarily concerned about the tooling, instead of the application.
Tooling level libraries are nice because they tend to be flexible enough for high data ink ratios, unlike highcharts, which turns me away with every example.
For large dataset and interactive visualization in Python, take a look at Chaco: http://code.enthought.com/chaco
R's power comes from the fact that hundreds of scientists have written packages for it when they have a new method - you won't be able to get that overnight. Also, R has strong links with other languages like C.
Finally, while I agree that sometimes R's syntax can be slightly obfusicated, I don't really think the examples on their site are fair... You can 'plot(y~x)' guys :p