Julia: A fresh approach to numerical computing
arxiv.org
arxiv.org
It makes very clever use of a really amazing new language feature in Julia 0.4, namely staged functions. At first it looked like it was manipulating C++ via strings sent to the Clang compiler at runtime, which sounded really slow and awful. But then it dawned on me how it really worked, and I was amazed.
I've actually convinced three or four of my colleagues (mathematicians) to work on building a computer algebra system with Julia that will interface to various packages such as Pari/GP, flint, Singular, possibly Gap. We meet weekly to discuss design and to try and work on an initial implementation.
The rate at which Julia is improving and the community is increasing is astonishing. We are at the stage where Julia is offering us language features we haven't even dreamed up a use for! And yet the language remains elegant and easy to learn.
stagedfunction foo(x)
if x == Int
:(2x)
else
:x
end
end
y = 1
foo(y) == 2
y = 1.
foo(y) == 1.0
Again, just like a macro, except that before `x` is quoted you work with the type of y instead of the symbol `y`. This means you can do even more crazy code specialisation, including on e.g. matrix dimensions or generating appropriate FFI calls based on arbitrary input types. It has zero overhead and you can use it just like a regular function.Very cool, and I think people will do really interesting things with this (especially after mixing them with macros).
See the wiki for a few bits of our planning that have actually made it online. Our current focus is twofold: 1) interface to Singular (http://www.singular.uni-kl.de/) from Julia 2) write a Singular interpreter in Julia (as an independent implementation of the Singular language). (Of course Julia will always be the main language of Nemo. The Singular interpreter will simply be a way for current users of Singular to benefit from Julia/Nemo and for users of Nemo to leverage the vast quantity of Singular library code out there written in the Singular language. And to have that code run faster of course.)
Yesterday we called the Singular C++ library, initialised it and created a Singular ring from within Julia for the first time. So very early days in that direction.
What is already committed on GitHub is a Julia interface to the flint library (which is pure C). Nemo is still a prototype, but you can Pkg.clone/Pkg.build it (Windows 32/64 support is there but clunky).
I will be visiting the Pari/GP people in January and working on a Nemo interface to Pari then.
I think this provides a very succinct summary of the Julia paradigm at a very high level.
This paper focusses on motivating the design of Julia for the larger scientific community. Any comments will be incredibly valuable in improving the quality of the paper.
Some ideas for improvements:
The package directory could be mentioned more prominently.
It's a bit tricky to select code in the PDF, not sure if much can be done about this?
Never in my life have I ever heard anyone say anything good about how wonderful it is language X is good at y which is what it was designed for. I have however heard plenty of people cursing languages for not doing something which they thought was an "obvious" thing for a language to do and X didn't.
Julia shines at developing efficient reusable code for crunching numbers. The type system is really nice, and the syntax is very clean and natural (to me!)
For some interesting ideas borrowed from Elm, do see: https://github.com/shashi/Patchwork.jl
Take for example the R programming language. Things you do in this language can be perfectly well done in e.g. Python too. Same goes for the MATLAB language.
Can you clarify which parts can easily be done in other languages and how?
Of course Lisp and Scheme can do this too, but it's virtually non-existent in any common language today. SICM makes heavy use of this.
Now, is there any other language with Lisp-like macros and static typing?
(: fun (natural natural -> real))
procedure signatures?edit: and _definitely not_ static typing.
I'm not sure if it is a very distinctive feature.
Note that I mean to symbolically calculate derivatives and integrals, not numerically. A homoiconic language can deconstruct any object representing code and can construct another object based on the underlying structure of the original, not based on any particular value of the original. The original doesn't even need to be invoked at all.
In mathematical parlance, a homoiconic language allows implementing functionals, as opposed to mere higher-oder functions (which is nothing more than function composition).
Please note that all this works with higher-order functions. You don't have to have a textual representation of the function. A.i. you don't parse your own source code.
LISP was actually invented to be able to implement functionals and symbolic calculations like this. The original paper demoed implementing function differentiation (among other things).
As far as Julia goes it aims to bring strong typing and very fast native performance for numerical operations. Neither of which R, Matlab, or Python provide. Seems perfectly reasonable to me.
[1]: http://en.m.wikipedia.org/wiki/S_(programming_language)
[2]: http://en.m.wikipedia.org/wiki/MATLAB
[3]: http://en.m.wikipedia.org/wiki/Python_(programming_language)
There's a direct and strong lineage, but they're not the same.
As someone who remembers when version 1.0 of R was released, I can say pretty confidently that it was well after 1976 :)
- It doesn't use the object.method() syntax, which is often more natural and readable than method(object).
- Support for interfaces but I'm not 100% sure about that.
It wouldn't really make sense for Julia to support object.method() any more than it would for Haskell to, superficial as it may seem.
Not all OOP languages look remotely like what you're describing.
You read mostly right. But the OOP language I like the most is Ceylon (it also has a first-class support for the functional paradigm).
I recently got into an argument with some coworkers over type classes. My argument was that it all made sense when viewed from the Dylan/CLOS/S4/etc lens, and that this was very much OOP, just not what they were used to. Many of them weren't buying what I was selling, but IMO they were using too strict a definition of OOP.
I didn't want it enough to flesh out the prototype (and would argue that Julia's native approach is better anyway) but it's inevitable that someone will do this eventually.
But look at the high-level, user APIs and you tend to see the more functional side – `fft(x)`, `plot(y)` etc., all pure functions which return expressions. Mutation is possible but discouraged via the `!` modifier.
Functions are the primary unit of abstraction – larger programs tend to be designed as collections of functions as opposed to object hierarchies. Higher-order functions like map and filter (aka "functionals") are encouraged.
So I tend to view Julia as a nice functional language which lets me drop down easily when I need to.
In terms of day-to-day programming, I'm not sure what makes Julia a better functional language than Python or Ruby. From my limited experience, it isn't speed.
I understand that work is under way to speed up anonymous function calls and to improve type inference on arrays resulting from `map`, etc. Once that is done then I will consider Julia a nice functional language.
That could get really confusing for the language's users, though, if there were name clashes between the two, and they would be there in any reasonably large program.
And of course, argument order already gives the first and last arguments special attention.
Another way to diminish that special-casing is going to the extreme other end: object.object.object.method(). Remove the now superfluous parentheses and replace the periods by spaces and you end up with Forth.
(object_1,...,object_n).method
I the former case u and v are of numeric types that fulfil some numeric properties, so no interfaces are necessary, they are implied.
I definitely agree with that, I said that object.method() style is often more natural. Typically when you have an object with some internal state and you want to send a message to it that will change its state. For example: threadPool.run(command) feels more natural than run(thread_pool, command).
You don't know that. I think that for people new to programming both styles can be natural, depending on the specific situation.
People discuss certain styles as being more natural, but in nearly every case (including this syntactical debate) I assert that it is solely due to what the most mainstream languages do and thus what people are more used to seeing.
That's not the same as being natural
https://en.wikipedia.org/wiki/Subject%E2%80%93object%E2%80%9...
But SVO is a close second, and certainly more common than verb-subject-object (the equivalent of `function(arg1, arg2)`).
From a mathematical standpoint what you are doing is to consider instead of objects X themselves, arrows from some fixed object I, x : I -> X (roughly speaking the "I points of X", in the category set I could for example be the one element set), then given another arrow f : X -> Y, the composition I -> X -> Y, would be denoted x . f, and you would actually find that in familiar cases this is precisely f(x).
However for this to actually be natural, ideally you wouldn't write x . add(y), but (x,y).add and if you follow that line of thought you would probably understand why stack based languages enjoy some popularity (the stack models the cartesian product).
Object oriented languages instead have for every I point x of X, a bifunctor hom_x(,), where hom_x(Y,Z) denotes all arrows from Y to Z, that "use" the point x (usually denoted by self or this). At least in mathematics, such an arrow can sometimes be extended to an arrow in Hom_X(Y,Z). In the case of addition above, you start off with + : (X , X ) -> X and use currying to get add : X -> hom(X,X) and pullback along x, to get $x^{\ast}add = add_x \in hom_x(X,X)$, which you then denote by x.add.
So you are right, the object method syntax is more natural and there even was a brief period in the 60s-70s when some mathematicians argued that it should be taught.
As a bootstrap to letting us do things like stream audio from the Web Audio API, and use Julia inside of an existing web app was to build a Node.js and Julia bridge:
https://github.com/waTeim/node-julia
It worked out very conveniently that Node.js through node-gyp can use C libraries, and Julia can run inside a C context.
I think the strongest use cases are node streams, and also online-learning models such as a recommender system. I.e. pass to Julia what the user has been interested in, and simply get back a few suggested items of high correlation.
I ran into problems with this when trying to implement an algorithm that creates and destroys lots of temporary objects, where each object is a wrapped C struct that can point to a large amount of extra memory (the algorithm in question is doing polynomial division, where each coefficient is a polynomial represented by a C object).
The Julia version of this program runs unexpectedly slowly and uses gigabytes of memory, even though a few megabytes should be enough. The same code runs much faster if you do it in C and free every temporary object when it's no longer used, or even you do it with a Python C extension. Since Python uses reference counting, the objects get freed as soon as they fall out of scope.
But Julia allocates thousands of objects at a time before running the GC. Inserting manual GC calls is even worse -- the memory usage drops, but the program slows down by an order of magnitude (every GC takes a long time, even if you've just allocated a couple of objects since the last time).
Part of the problem is that the C code in question uses an object pool to speed up frequent allocations and deallocations. But even if you disable the pool, Julia performs worse than you would hope. A good generational/incremental GC would solve the problem.
https://github.com/JuliaLang/julia/pull/8699 https://github.com/JuliaLang/julia/pull/5227
* Julia's approach to OOP is via multimethods, not the usual class/inheritance model. This might be annoying for people who don't want to learn how to be an effective programmer in the other paradigm.
* Julia's garbage collector is not generational/incremental and in some corner cases, GC can take 10 times longer than the actual function you are running (typically it is between 5% and 50%). This would make Julia unsuitable for HFT, real-time games, web browsers of the future and other real-time applications. (Edit: see Viral's post in the same thread. Incremental GC is in the works. This is genuinely my experience of Julia to date. No sooner do you need something, and someone competent is already working on it, if they haven't already done it!)
* Julia sucks at predicate dispatch. It doesn't have it. Granted, neither does any other language except Gap and one other I forgot. So if you are used to that feature, you would find Julia a step down. (Edit: yes of course Julia does not need/want predicate dispatch. It's just an illustration of something that could bother you if you were really, really used to something. I just happen to have colleagues who really are used to this.)
* Julia is tied to LLVM, so if you want to be on the CLR/DLR or JVM, you are out of luck.
* Julia currently doesn't have static compilation (I hear it is being actively worked on). This makes it more difficult to deploy binaries.
* Julia does not have Haskell-like separation of effects from pure functions. This might not appeal to type purists.
* Julia functions can fail at runtime where statically typed/compiled languages would pick up the errors at compile time.
* As popular as it is, Julia is still not in the top 50 programming languages by measure of usership.
* Julia's support of mutable C-style structs allocated on the stack, as opposed to pointers to heap allocated objects is still somewhat lacking. This creates some challenges in efficient C FFI in corner cases, especially in combination with GC.
* Julia doesn't have inheritance of data types (it is really a dual to a data focused language, which is sensible -- you typically have far more functions than data types in a program -- but it's still hard for traditional OOP users to get used to).
* The Julia abstract type system is somewhat linear, which makes it a little less flexible as far as contracts/interfaces are concerned, if you choose to implement things that way.
Of course Julia has so many features that make it worthwhile, that it is worth investigating for many projects. It has multimethods, (static) dependent typing, very easy and efficient C interface (soon C++ interface), C-like performance is possible, garbage-collection, macros, runtime console (REPL), great numerical features, Jit compilation, a good selection of libraries/packages, a package manager, profiling, various development tools, a good (highly intelligent and helpful) community.
Back to Julia: lack of CUDA support was a limiting factor for me when I tried it a year ago.
However, we personally feel that the path Intel is taking with Knight's Landing, is likely to be the best of both worlds - GPU and general purpose - assuming it materializes for real in 2015.
* module compilation is not cached, so if you use very many modules your start-up times can be slow
* error messages sometimes require some head scratching. For example, it's not uncommon to get an error that there's no available function convert(::SomeType, (Some, Args)) when there's no obvious convert() to be seen in the code in question. Occasionally the stack trace will be missing from errors, or there won't be a line number. Obviously this is improving quickly, but can be frustrating.
https://github.com/JuliaLang/julia/pull/8656 https://github.com/JuliaLang/julia/pull/8745
Being tied to LLVM is real, but we do have JavaCall.jl to call Java code.
https://groups.google.com/forum/#!topic/julia-users/hxfR70Ro...
I also find the dataframe library to be much harder to use/not much faster than pandas.
Great community though.
I think that being not faster than Pandas is pretty good, since it is written in C and heavily optimized.
1-based indexing is very common in mathematical software.
I'd much rather have +1 scattered throughout my code (though I tend not to) than have to do -1 in my head at the repl, personally.
In R computing an outer product while passing a custom function is easy. Something like outer(x,y, some function of (x,y)) is pretty straight forward. Not sure how to achieve this in Julia.
Nevertheless, as the language grows and as more idioms are documented and widely used, lif will be simpler I guess.
http://www.meetup.com/JuliaChicago/events/216950712/
Should be a great way to get introduced to the language and pose your nagging questions to a community expert.
I have been weary of the Julia pitch which was basically "lets brew R, Python, and Matlab (all >20 years old with massive and unrivalled ecosystems and all serving their purposes very well, thank you) into one "new" language and put a nice logo on it". How is this language seriously the leap forward that one needs to abandon the current awesome toolsets, with all their battle tested libraries, other than a nebulous "speed" argument which in many cases, judging by the comments and my own experience in financial data matrix operations, is not even fully accurate? Ready to be persuaded otherwise if I can be shown that other than "speed" there are hefty reasons to move from Python and abandon the almost endless choice of richly varied tools that I have at my disposal already.
Then there's the really powerful metaprogramming capabilities, which are absent from the languages you've mentioned. I can't emphasise it enough: the paper this comment thread is about does a great job of explaining why these things are compelling.
Also, you may be interested in [1], an argument that language performance is valuable even if you don't need it yourself.
[1]: https://medium.com/the-julia-language/performance-matters-mo...
I should add one more point though. The post mentions Matlab 15 times and Python only 7. It's possible this whole Julia effort will be successful with the Matlab crowd which, up to now, has been watching with horrified fascination from the sidelines as open source ate its lunch.
I really want to love Julia and everything I read about it makes me think it's the language for me, but every time I try to write non-toy code the performance is always worse than python.
edit: Nevermind! On rereading the mailing list thread (especially this message [1]), I might have mischaracterized the issue slightly. It probably came from speed issues with small integer powers (i.e., there are faster ways to exponentiate small integer powers than there are for general exponentiation that Julia hadn't been using but other software like matlab and octave were), and that issue in Julia seems to be fixed now.[2]
[1] https://groups.google.com/d/msg/julia-users/n3LfteWJAd4/jag6...
[1] http://julia.readthedocs.org/en/latest/manual/performance-ti...
edit: As a direct reply to your question, yes, lots of people fail to find performance improvements when they first port code to Julia. But they often discover those improvements 10 minutes after someone with more Julia experience looks over their code :)
Then it is not surprising at all that you do not get any speed benefit. I would expect a serious slowdown.
Idiomatic MATLAB and Numpy has lots of vectorized expressions because they are so crappy at loops. However, in vanilla Julia those vectorized expressions cannot be JIT'ed. Now if you were to write them out as explicit loops then JIT will take a stab at making things faster.
I do like the economy of words of vectorized expressions a whole lot, this is the reason why I am so excited about https://github.com/lindahua/Devectorize.jl
My goal is to port a lot of my code to set up astrodynamics simulations. I've even created a placeholder Github repo for it [1]. Just have to find the right strategy to involve the dev in my existing projects and get my colleagues to chip in.
Anyone interested in using Julia for bioinformatics is invited to contribute to, or follow the progress of, the BioJulia project: https://github.com/BioJulia/Bio.jl
Winston is a most traditional plotting API:
http://winston.readthedocs.org/en/latest/examples.html
You can also plot seamlessly via Python using PyPlot:
https://github.com/stevengj/PyPlot.jl
They are all quite well documented.
Perhaps it's all very obvious if you're already familiar with ggplot2, but as somebody without an R background, I found their docs extremely spartan.
In the rare cases of a surge, such as an HN mention, it will auto-release account approvals as the surge subsides in a few hours. Once you do have access, you will continue to have access, even during surges. Basically, we are just trying to manage our compute budget better.