The Julia Programming Language
julialang.org
julialang.org
Personally I'm a little wary of being ghettoised into something overly domain-specific for scientific/numerical computing. Really good interop may mitigate that -- something which can navigate the unholy mix of C, C++, fortran, matlab, octave, R and python routines one comes across trying to reproduce others research work, would indeed be awesome.
I do wonder if some of the noble demands of this project might be better delegated to library developers though, after adding a bare minimum of syntax and feature support to a powerful general-purpose language. For now Python+numpy+scipy seems a great 90% solution here.
From the article: "The library, mostly written in Julia itself, also integrates mature, best-of-breed C and Fortran libraries for linear algebra, random number generation, FFTs, and string processing."
However, much of the code is written in a vectorized style for performance reasons, as is the case with many high level scientific computing languages. This leads to unnatural code in some cases, and also uses too much memory. First I thought IronPython was the way, but have been looking forward to pypy+numpy+scipy eagerly.
If I were to use julia, the code would be a lot more natural, because type inference and all the compiler goodies make it possible to simply write loops over arrays when I need them. This was one of the reasons we started working on julia, because everything else just seemed to fall just a little bit short.
The parallelisation story isn't great at the moment either, although seems like it has the potential to improve. Still 'sucks' is pretty harsh. For me, looking at how far it's come since I first played with it, I'm impressed. For machine learning, the majority of the building blocks one needs are there, and you get to sit back and put them together using a nice, clean, widely-adopted general purpose programming language. And unlike MATLAB still maintain a decent amount of control over things like memory usage and which BLAS routines it's calling.
Adding bindings for new libraries is more of a pain than it should be though on the occasions where you do really need some fortran or C++ library that doesn't have bindings yet. A language which bridges the gap between high and low levels (not C++!) and has great interop would be very interesting. I guess I'm just hopeful that this kind of thing can be achieved in a general-purpose language (new or existing) which like Python is adopted across the wider software engineering community. Perhaps that's my unreasonable demand to add to their list :)
Seriously? I mean, I know you are very excited about your work on PyPy, but why do you have to go around saying baseless stuff like this? There are many, many thousands of people who write lots of Python code every day that rely on Numpy, Scipy, and Matplotlib to get their work done, and who seem to be quite pleased with it. There is a lot of work left to do on all parts of that stack, but that's a far cry from "it sucks".
The general consensus, if you look at languages like Chapel and Fortress and X10, seems to be that most scientific codes shouldn't be written using for-loops. That is the low-level control flow construct dating from the age of assembler. Instead, what scientists generally want to say is, "Apply this kernel across this domain, with these windowing conditions", or "Reduce values from this computation along these keys in my dataset". As software developers, our job is to provide the language runtime to allow them to do that; such a runtime will be the most robust, correct, maintainable, and performant.
The problem with numpy's performance is twofold:
* Numpy expressions might not be fast enough. I believe you guys at continuum are trying to address that one way or another. In general the kernel expressed using high-level constructs in python should not be slower than an equivalent loop in C.
* Sometimes you actually want to write a for loop, because you don't care, because it's faster, because it's a single run, because the data is manageable etc. You should not be punished for doing that with 100x performance drop. You can still be punished for that with 2x performance drop.
From a quick look I can tell that while Fortress is more similar to Scala, with classes, mixins and static typing, Julia is closer to Clojure, with no encapsulation, dynamic typing, homoiconicity, and separation of behavior (methods) from concrete types (akin to Clojure's protocols).
I really like the choices Julia's designers have made, particularly multiple-dispatch methods, final concrete types and lispy macros. The language seems very elegant. Not too crazy about "begin" and "end" syntax, though :)
Chapel has made a lot of progress in the past 2.5 years (while we've been working on Julia) and is certainly a contender.
Julia certainly aims to be more dynamic like Clojure, but specially designed to be good for numerical and technical stuff, and yes, the begin/end is due to Matlab — scientists who have a lot of code in Matlab already are a major target demographic. As it is, a lot of Matlab code ports over with relatively trivial changes (see: http://julialang.org/manual/getting-started/#Major+Differenc...), which was the point of syntactic similarity to Matlab.
Octave has solved this by allowing both matlab-style "end", as well as block-specific endings, like "endif", "endfunction", "endfor", etc.
Having several end end end endings without any hint of what they are closing is one of the ugliest parts of matlab. Even though you are reimplementing this for compatibility reasons, it is rather sad that you chose to make this the default behaviour for Julia.
If I weren't excited about it, I wouldnt even be commenting on it. ;)
I disagree. Besides (), [], {}, and <>, there's plenty of other paired punctuation available in Unicode, such as...
⦃⦄⦅⦆⦇⦈⦉⦊⦋⦌⦍⦎⦏⦐⦑⦒⦓⦔⦖⦕⦗⦘⧘⧙⧚⧛⧼⧽〈〉《》「」『』【】〔〕〖〗〘〙〚〛❨❩❪❫❬❭❮❯❰❱❲❳❴❵⟦⟧⟨⟩⟪⟫
Nothing is as bad, though, as the default Matlab behavior of not indenting first-level function code. This makes it really hard to scan a file with multiple functions and see where they separate.
When I said "easy fix" I meant "a trivial extension to the parser" without realizing you would translate that to "really low priority". Guess I'll save my non-trivial suggestions. :)
Braces would be better than ending keywords, in my opinion, but I'm disappointed not to see whitespace-defined blocks of a la Python. Perhaps that doesn't work in a statically-typed, type-inferred language, but if it does, I think it would be much better. Scientists have no problem with it (certainly less of a problem than programmers do) and Python is pretty well accepted in the scientific computing community.
http://bartoszmilewski.wordpress.com/2011/11/07/supercomputi...
After that proposal is implemented (perhaps in version 3 of Scala, almost ten years after its first release, though granted it wasn't their top priority) they'll still have to write at least a little bit of boilerplate for each type they want to store in unboxed arrays, and they'll be limited to types that can be stored as Java primitive types under the covers. That seems like a pretty nasty limitation for a general-purpose scientific computing language, not only because it decreases native performance but because it makes C interop much more difficult.
For a general-purpose scientific language, I think it is imperative to adopt the principle that any user-designed type should be able to achieve the same power, performance, and elegance as if it were a built-in type. To the best of my knowledge, that rules out JVM languages.
(Scala was never meant to be a "JVM language," but in terms of its strengths, its weaknesses, its user base, and its future, it is very much a JVM language, and the obstacles to implementing efficient user-defined types is one of the trade-offs they accepted when they went with the JVM.)
http://ppl.stanford.edu/main/publications.html
example, this scala DSL:
So interop is a very, very high priority. We have pretty good C ABI interop now. You can just call a C function like this:
ccall(:getpid, Uint32, ())
and it works. It works in the repl too and you can dynamically load new libraries in the repl. See this manual page for more details: http://julialang.org/manual/calling-c-and-fortran-code/.We still want better C/C++ interop though. I've already looked into using libclang so that if you have the headers you don't even have to declare what the interface to a function is. That was very preliminary work, but I got some stuff working. Making Julia versions of C struct types transparently would be another goal here.
Another major interop issue is support for arrays of inline structs (as compared to arrays pointers to heap-allocated structs). C can of course do this, but in any language where objects are boxed, it becomes very tricky. We're working on it, however, anyone who wants to discuss, hop on julia-dev@googlegroups.com :-)
Oh, so arrays of complex numbers are implemented as arrays of pointers to objects? That would give really bad CPU cache performance.
The longer-term approach is still up in the air and that's what I was talking about above. My favorite approach at this point is to allow fields to be declared as const — which in Julia means write-once. Then if all fields are const the object is immutable and can automatically be stored in arrays inline.
What's the problem to have arrays of doubles and complexes as "basic" types even in the dynamic language? I believe this could give you a C footprint and C performance with array operations?
Naive question (and I'll hold my tongue on my naive guesses): I understand why it is necessary to box user-defined types on the JVM, but why build this restriction into a new language that doesn't run on a restricted platform? Especially when performance and C interop are high priorities?
Or perhaps I misunderstood, and your statement should be read as sparsevector suggested.
It seems that in dynamically typed languages, you either need to have two kinds of objects (value types vs. object types), or two kinds of storage slots (e.g. arrays that hold values inline vs. arrays that hold references to heap-allocated values). The boxing aspect is really only part of that since you can't get the shared-reference behavior unless the storage is heap-allocated, regardless of whether there's a box header or not.
So yeah, it's a complicated issue.
One other slight difficulty seems to be the lack of finilisers, e.g. to clean up a GMP bignum when there are no remaining references to it.
The fact that it has parametric types, parametric polymorphism, macros, performance almost as good as C, good C/Fortran interop, 64 bit integers and an interactive REPL all in the one language just blows my mind.
I wasn't able to tell if it is possible to overload operators, which is another thing essential to mathematical code.
I was also unsure why the keyword end was needed at the end of code blocks. It seems that indentation could take care of that.
I also didn't see bignums as a default type (though you can use an external library to get them).
However, all in all, I think this is the first 21st Century language and find it very exciting!
https://github.com/JuliaLang/julia/blob/master/j/array.j#L15...
I personally would find multiple indexing schemes confusing both for usage as well as to develop and maintain. Given that 1 based indexing seems to be a popular choice among many similar languages, we just went ahead with that.
It generalizes the choice of 0 or 1 to an arbitrary starting index. So when you create an array you specify not just where it ends but also where it begins. This lets you do neat things (consider a filter kernel with range [-s,+s]^n instead of [1,2s]^n) and the extra complexity it adds can be hidden when not needed using for-statements or higher order functions.
Nobody uses it because the implementation is not very efficient and Haskellers have a chip on their shoulder about performance. It subtracts the origin and computes strides on every index, but you could easily avoid this by storing the subtracted base pointer and strides with the array. Of course when you go to implement it you'll see light on 0-based indexing :)
The only way I could think to manage it would be a pragma which switches 0-based on for a given file. But this is doubtlessly not trivial.
Or there is the less elegant option of introducing a new operator for 0-based arrays. But this is liable to cause confusion I think.
Also see Dijkstra's take on the matter: http://www.cs.utexas.edu/users/EWD/ewd08xx/EWD831.PDF for
However, as I've used it more and more, 1-based indexing has really grown on me. I feel like I make far fewer off-by-one errors and actually hardly ever have to think about them. This has led me to conclude that 1-based indexing is probably easier for humans while 0-based indexing is clearly easier for computers.
Be careful when generalizing from your own personal preferences and cognitive biases to what is easier for humans in general.
I find 1-based indexing to be weird/illogical and prone to off-by-1 errors.
Inclusive/exclusive ranges as in Python along with 0-based indexing means the likelihood of any off-by-1 is negligible...
Lua also has 1-based arrays. Heh.
The language syntax seemed uninteresting, which is not really a bad thing but what's the case for the need for a whole 'nother language in that category?
I don't think lua has integers ie. like Javascript everything is a double, that can be changed by redefining a macro and recompiling AFAIK but it's still one global numeric type (and you can't change it for luajit).
I have been working on and off on a BSD licensed bignum library. Maybe in a few years....
Also, LLVM handles carries and certain loop optimisations poorly, so even using LLVM bytecode you can't do much better than compiled C. It would be a massive project to improve this in LLVM (I thought about giving it a go sone time ago but decided it was overwhelming). And that use case is probably too specialised for the improvements to help with much else. Obviously the LLVM backend is fantastic for 99% of use cases and improving all the time.
N.B. I am not implying that a good assembly programmer is generically faster than a C compiler. Bignums are a very special case.
Various libraries used by the Julia environment include their own licenses such as the GPL, LGPL, and BSD (therefore the environment, which consists of the language, user interfaces, and libraries, is under the GPL). Core functionality is included in a shared library, so users can easily and legally combine Julia with their own C/Fortran code or proprietary third-party libraries.
The FSF disagrees (http://www.gnu.org/licenses/gpl-faq.html#NFUseGPLPlugins):
If the program dynamically links plug-ins, and they make function calls to each other and share data structures, we believe they form a single program, which must be treated as an extension of both the main program and the plug-ins. In order to use the GPL-covered plug-ins, the main program must be released under the GPL or a GPL-compatible free software license, and that the terms of the GPL must be followed when the main program is distributed for use with these plug-ins.
1. The Julia core, which consists of the language runtime and core functionality, is MIT licensed, and builds into an MIT-licensed shared library.
2. The Julia "environment", which includes a user interface, third-party libraries, etc., some of which are GPL, and which is therefore GPL as a whole.
I believe they're saying that you can link #1 with proprietary code, not meaning to imply that you can link #2 with proprietary code (because as you point out that wouldn't work). How useful that is probably depends on how many of the libraries the average application needs are in bucket #1.
Given how similar to python they are their main competitor is pypy, which isn't too far behind them in performance I'd suspect (absent from the benchmarks though). I was fully expecting someone to make a non-compatible fork of pypy to create a faster python-like lang/something like julia. They should quickly get some elegant small libraries like sinatra/bottle and requests (access to dbs with http interfaces as well). A robust requests+lxml library built-in would be simply amazing.
Further, if we are going to be nitpicky, when you use something like that for its intended purpose it's still a fork - I forked Twitter bootstrap for my new project/I used Twitter bootstrap for my new project - so "you wouldn't fork PyPy" is a false statement.
> if we are going to be nitpicky, when you use something
> like that for its intended purpose it's still a fork
No. Fork is definitely not semantically congruent with use. You're only forking Bootstrap if you create a path that diverges from its mainline; don't get confused by GitHub parlance.There seems to be a disease, let's call it Eric Raymond Disease (ERD), on HN in particular[2] of people busting in rudely to declare with supreme confidence on the usage of quite generic words and/or slang. Let's start by tabling the fact that a particular "software definition" of "fork" has not yet reached the OED, or online dictionaries like Merriam-Webster for that matter. Going from that I can't see how you can declare it to only have one definition when you already admit there is a competing one ("GitHub parlance"). The base concept of the word fork is a divergence, a branch, which seems to me to cover any copy+modify move for software so someone copy+modifying bootstrap is forking it, someone copy+modifying the implementation of python called pypy is forking it.
Are you and Daslch really contributing to HN or are you taking a rather banal comment and turned it into a 100% useless thread by bickering over semantics?
[1]http://pypy.org/ [2]http://news.ycombinator.com/item?id=3467035
fib - roughly the same time as CPython
parse_int - 10x over CPython which is extrapolating to their python vs julia about 20% slower
quicksort - 4x faster than cpython much slower than julia.
pi_sum - 20x faster than cpython, a bit faster than julia?
Those are very unscientific measurments. Both fib and qsort are recursive benchmarks (why would you write qsort recursively??), so I guess the JIT does not have time to kick in (and PyPys support kinda sucks for recursion at least in terms of warmup times).On a slightly unrelated note - a modified lxml library runs on pypy albeit a lot slower right now.
Cheers, fijal
EDIT: updated formatting
Unfortunately, this one combines the inconvenience of Template Haskell (explicit invocation of macros with @) and the bugs of Common Lisp (the section on hygiene says, basically, "we have gensym"). Fortunately, the language is young, and I hope they can improve this story.
I find that the @ syntax is amazing, since the reader (programmer) knows exactly which form is a macro (and can look it up), as opposed to introducing arbitrary, often very confusing syntax.
So I'm wondering does a programming language which marks macro expansions different to function calls (as Julia does with the @ prefix) really need to distinguish functions from macros in the definition (as Julia does with the 'function' and 'macro' keywords) ?
C or fortran can be as fast as they want, but in mathematica I can do MorphologicalComponents[] and get the components. having these functions available to me speeds up my time to discovery by 1000x or more.
I came across this bit of humor a while back that communicates the general feeling well:
http://bit.csc.lsu.edu/~gb/csc4101/Reading/gigo-1997-04.html
If the problem IS that it's huge and inelegant, then the solution isn't to try to improve it, but to start from a clean slate. Julia looks like a reasonable attempt.
import time
def fib(n):
if n < 1:
return None
if n == 1:
return 0
if n == 2:
return 1
x0 = 0
x1 = 1
for m in range(2,n):
x = x0 + x1
x0 = x1
x1 = x
return x
if __name__=="__main__":
assert fib(21) == 6765
tmin = float('inf')
for i in xrange(5):
t = time.time()
f = fib(20)
t = time.time()-t
if t < tmin: tmin = t
print str(tmin*1000)The story the micro-benchmark tells is that Julia's pretty good at function calls, but JavaScript is even better (and, of course, C/C++ is the gold standard). In general the V8 engine is really amazing. We had the advantage of being able to design the language to make the execution fast (with the constraints of all that we wanted to be able to do with it), but V8 makes a language that was in no way designed for performance and makes it blazingly fast.
After I wrote the benchmark code for JavaScript and saw just how fast it was I had a moment of "should we be doing scientific computing in JavaScript?" Now wouldn't that be nuts?
one notable restriction is that inheritance is only for interface, not implementation.
also, can anyone find a sequence abstraction (like lists)? arrays seem to be fixed size and i don't see anything else apart from coroutines. am i missing something?!
[perhaps not, if it's intended for numerical work. on reflection i am moving more and more towards generators (in python) and iterables (in java, using the guava iterables library to construct maps, filters etc) rather than variable length collections, so maybe this is not such a big deal. it's effectively how clojure operates, too...]
Not quite. You do inherit methods defined on abstract types.
Arrays are not fixed size: there are push, pop, shift and unshift operations on them just like Perl, Ruby, etc. This uses the usual allocation-doubling approach so that the entire array doesn't need to be copied every time, but it's still usually much better to pre-allocate the correct size. Of course for small arrays that are typical in Perl, Ruby, etc., it hardly matters. If you're building an vector of 1 billion floats, however, you don't want to grow it incrementally.
The sequence/iterable abstraction is duck-typed: an object has to implement methods for three generic functions:
i = start(x)
done(x,i)
next(x,i)
The state i can be anything.Lack of implementation inheritance is one of those things that intro to OO books make a big deal of, but when you don't have it, you don't miss it at all — or at least I don't. I've never found tacking a few fields onto the end of an object to be very useful. I don't want to inherit memory layout — I want to inherit behavior. Julia's type system lets you write behavior to abstract types and inherit that for various potentially completely different underlying implementations.
Must compound types in Julia be concrete? I don't think I saw it in the documentation.
So far we haven't actually felt a "pressing need" for multiple inheritance or interfaces, and we tend to take a pressing-need approach to language features. If you can live without a feature for a while, then maybe you really didn't need it in the first place. But we'll have to see what happens when other people are starting to try to use if for things.
Aren't compound types inherently concrete? The compoundness describes the implementation of the type, implying that it must have an implementation, hence must be concrete.
brew install wget
[Edit] git page says gfortran (and wget) are downloaded and compiled, but if they're not already installed make fails. So...
brew install gfortran
The need to do this separately may have to do with licensing?
[Edit] And if you're not root, install to /usr/share/julia will fail. So you'll need to do:
sudo make install
I'm sure all this is perfectly obvious to Unix-heads who are inured to this sort of abuse, but I'm a Mac-head, used to things that Just Work, and I hate this shit.
Stepping back a bit, this is one of the reasons why having an entirely web-based experience is appealing — then you can let people use a known-good setup without needing to mess around with installing a fairly extensive amount of software just to get basic things to work. Then there's also the general appeal of doing big data work a la Google docs or Gmail. The trick is getting the user experience to be good enough on the web.
I once designed/implemented a language where the biggest mistake I made was using 1-based indexing. At first it looked like it was going to be easier to understand/use but 0-based indexing is actually much more convenient when dealing with indexes arithmetic.
PS: the "metaprogramming" stuff is fabulous.
The webcast is delayed and there's always something interesting that happens after the camera goes off.
Here is the certificate error: Connecting to github.com (github.com)|207.97.227.239|:443... connected. ERROR: The certificate of `github.com' is not trusted. ERROR: The certificate of `github.com' hasn't got a known issuer.
BTW, do try out the mac binaries if the build is an issue. We are still trying to make it all build seamlessly!
http://loopkid.net/articles/2011/09/20/ssl-certificate-error...
Basically, you need to get the right certs and then tell wget to use them.
One more thing: it would be awesome if the `manipulate` from Mathematica could be incorporated somehow in that web interface. See:
I do believe that open source + good compiler + simple C/Fortran calling interface will lead to others being able to write toolboxes in julia itself and plugin libraries when needed.
http://technicaldiscovery.blogspot.com/2011/10/thoughts-on-p...
Not only do some of the example's seem to be created specially to show the benefit of the JIT compiler (see pisum fe). Most of the example's do not really use the features of LaPack/BLAS, the one where it does (a matrix multiplication in rand_mat_mult) shows that all languages which use these optimized libraries beat Javascript and handwritten C++ with a magnitude of 2.
Thus, be careful with making such a generalisation from this benchmark. Also, it is much simpler to simply work with a language with proper support for multi-dimensional matrix slicing than having to do this all by hand.
Julia does have multi-dimensional arrays and slicing also. If you had to do all this by hand, it would just be simpler to use C. http://julialang.org/manual/arrays/
One of the key concepts behind Julia is that not only the end-user, but also the numerical library writer, should benefit from using a high-level language. In Julia, almost all of the library code is in Julia, and it's as fast as the library code written for R, Matlab or NumPy in C. There are also a lot of situations where vectorized code is either awkward or inefficient — especially in terms of creating a lot of unnecessary temporary arrays. Languages where the high-level language is slow force you to do everything vectorized — in Julia, you're not forced to do that. If you want to write a C-style scalar loop, you can and it will be fast (and you don't even need type annotations to make it fast, as shown by the benchmarks).
V8 is really impressively fast, but JavaScript as a language is not very well suited to scientific or technical computing.
Same in Programming Language design. Few original concepts, Lisp,C,Smalltalk, but so many cocktails even my grandma is creating one. All we want is more libraries.
I see that d3.js is included in the package, but a cursory glance at the docs and I didn't see any examples of how to generate a chart.
MathKernel -script "foo.m" -noprompt
As to Mathematica being a pain to write code for: I've found it to be a pretty nice language, certainly a far cry better designed than MATLAB.Saving to: `lapack-3.4.0.tgz'
2012-02-19 01:03:26 (55.3 KB/s) - `lapack-3.4.0.tgz' saved [6127787/6127787]
make[3]: gfortran: Command not found make[3]: * [lsame.o] Error 127 /bin/sh: ./testlsame: not found /bin/sh: ./testslamch: not found /bin/sh: ./testdlamch: not found /bin/sh: ./testsecond: not found /bin/sh: ./testdsecnd: not found /bin/sh: ./testieee: not found /bin/sh: ./testversion: not found make[2]: * [lapack_install] Error 127 make[1]: * [lapack-3.4.0/INSTALL/dlamch.o] Error 2 make: * [julia-release] Error 2
Great work guys.
JULIA sys0.ji Segmentation fault make: [sys0.ji] Fehler 139
(also, julia is not a great name for googling....)
The R thread on the dev list gives an accurate representation of the obstacles to Julia being widely adopted, however. One might become very unhappy while using R, but sometimes you have to use it, because something you want is only in R, due to the ubiquity of that tool for stats.
If Python, which does not have that many disadvantages besides not having been explicitly designed to appeal to users like me, cannot unseat Matlab and R, Julia will have a difficult time.
But I will give it a try, and will implement some simple but core algorithms that I use a lot.
Another idea to test Julia's usefulness would be to port a tool like Waffles to Julia. In my opinion implementing such a sensible tool like Waffles in C++ is a heartbreak.
Sound interesting though if it can make hadoop easier to use id take that as win I dont think its going to replace fortran as a HPC language.
Julia's syntax is intended to be familiar to users of MATLAB®. for i = 1:t
in Fortran: DO I=1,T
The ability to load both C and Fortran dynamic libraries is also really really nice.Does 0-based indexing have any advantage, beyond ability to do pointer arithmetic with indexes?
To be honest, if you're targeting early adopters of a programming language Linux and Mac support is probably a lot more important than Windows. Smart Windows users can always run Linux in a VM.
Please, be a minimum objective and constructive.
It is not by design that we do not have Windows support. It's just that none of us uses windows. We do believe that code is largely portable, and with a little effort, it can be built on cygwin or mingw. Nothing like a native port though. Maybe someone who is familiar with windows will come along and contribute. This is a common question our friends and colleagues ask us.
That doesn't sound trivial to port to Windows, and I'm sure their own application was their first priority.
It's about standards, but it still applies.
I'm sure that it's filling a real need. I just wish that groups could cooperate to make a handful of languages/libraries better rather than having 100 competing ones.
The different language approaches will typically compete on things like syntax, speed, etc., and this leads to innovation and cross-pollination of ideas. I personally prefer to have some choice, and use a number of languages for scientific computing myself - but I guess too much choice makes things confusing for the newcomer.