Linear Regression 3 Ways in Julia
boss-level.com
boss-level.com
The following blog post goes into the reasons, and the remedies. In particular, there are ways to automatically devectorise your code.
If you mean loop merging, the Repa 4 library does something pretty interesting called series fusion. But it also needs a compiler plugin so that it can do that optimization internally.
I'm actually working on my own vectorization / fusion friendly numerical tools in Haskell. One philosophical point that's pretty darn important Is making it obvious and clear how to get good performance in a robust way. Many auto vectorization codes (eg try auto vectorization in clang) require a bit of work to tickle correctly. So in many ways, I think the best approach is to encourage idioms that don't require compiler smarts.
http://ghc.haskell.org/trac/ghc/ticket/915
It's still one of the most impressive moves to a sufficiently smart compiler.
The different fusion strategies have different trade offs, so it could also be that I'm not reading your with the correct nuance.
At the moment, SF seems to support natural looking code, but not always truly natural code. It's tantalisingly close, but it doesn't seem to be getting any closer.
On the other hand, I've seen (and written) plenty of code that went through contortions to use vector operations in IDL (http://en.wikipedia.org/wiki/IDL_(programming_language)). It would be great to have a language where for loops and vector expressions are fast.
The real point of Dahua's "Fast Numeric" blog post [1] is that you can devectorize in the high-level language if you need to, not that you always should. Most code is not a performance bottleneck and can use the most convenient form. In cases where you identify a real hotspot [2], in Matlab you'd have to resort to writing C code via MEX. In Julia, you do one or more of the following: (a) change some copying function to a mutating version by sticking a "!" on the end of it; (b) move an array allocation out of a loop and pass the pre-allocated array into a mutating function inside the loop; (c) manually devectorize some code if absolutely necessary – or use the Devectorize package to do it automatically. At no point do you ever have to resort to writing code in any other language.
Also keep in mind that slowness is not the only problem one encounters with numerical code – I've personally have much more frequent trouble with memory usage. I've written far too much insanely convoluted Matlab code to try to avoid blowing out the memory on a 100GB machine (or 1TB these days). It's generally just barely doable using find, sparse, sub2ind and ind2sub. If I had to resort to writing C code every time I encountered one of these memory problems, I would just write the damned thing in C and scrap Matlab altogether. On the other hand in Julia for such situations, I can just write a few for loops and not only do I avoid blowing out my memory, but the code gets faster too, and I'm never required to leave my high-level environment.
https://groups.google.com/forum/#!topic/julia-users/T5YYKQzl...
Σ₁,Σ₀, σ²
rms = √(1/len(x)*Σ(x))
Which is really quite cool from a mathematical view point.