Julia Computing Granted $600k by Moore Foundation
moore.org
moore.org
I guess Python is "the competition" so here's my thoughts on that competition (from knowing fairly little about Julia). From my limited experience Python is actually easy enough to grasp for most scientists without a programming background so I'm not sure the "easier" argument is what should be focused on. At the "lower end of the spectrum" (that is not super hard data crunching and/or less affinity for programming) there's social sciences, psychology, marketing, economics etc. and they use a variety of higher level tools (the pesky SPSS) for their routine tasks and are usually willing to invest time for more data crunchy activities (often R but Python is common as well in say economics with data-crunchy models).
The big attraction of Python (for me) is that you don't just get the data crunching but also the boilerplate you'll need around it (data normalization, stroing stuff in databases and getting it out, turning your models into simple web-apps etc.) + IPython notebook and distributions like anaconda make setup and quick experimention fairly simple.
If I were to give advice to someone entering science I'd certainly suggest "learn python". However...the more languages we have the better, MATLAB makes me cringe for multiple reasons. Go Julia :D
https://blog.jupyter.org/2015/04/15/the-big-split/
I think Julia was one of the early alternate languages.
http://blog.dominodatalab.com/lesser-known-ways-of-using-not...
http://ipython.readthedocs.org/en/stable/interactive/magics....
See e.g. https://github.com/DELabUW/szkola-letnia-2015/blob/master/za...
Python fans are a bit unnerving, they seem to be more vocal among the similar alternatives R, Octave/Matlab, SPSS. Some people simply prefer Julia for its nice language syntax and its native raw speed, deal with it.
It'd be difficult to list four languages that are more different than those in your list. There's some overlap in the tasks each is suited for, but they're hardly alternatives.
I didn't want to get into this holywar, but god darn it, what kind of half-assed language uses the notation C = A.dot(B) for C = A * B? It's as if the authors of Numpy took just enough classes of Linear Algebra 101 to learn about dot products, and then dropped out of college to pursue their glamorous careers developing their Matlab knockoff.
As a long time Python user, I disagree with the insinuation of competition between the two communities. In fact, I've met all of the core developers of Julia and of the major Python data science projects, and they see each other as fellows and colleagues!
For example, last year at PyData NYC, Stefan Karpinski gave the following talk, titled 'Julia + Python = \N{heavy black heart}'
https://www.youtube.com/watch?v=PsjANO10KgM
This past week, at PyData NYC, Stefan Karpinksi also shared the stage with Andy Müller (scikit-learn) and Jared Lander (author of R for everyone) to discuss the future of machine learning with the audience and addressed how the various tools, languages and communities fit together.
In fact, most people may not realise, but Julia, Jupyter, IPython, numpy, pandas, matplotlib, &c. are all fiscally sponsored projects of the same 501(c)3 non-profit, NumFOCUS (which is the organisation that also runs the PyData conference series)!
As a person who has spent the best part of year coding almost exclusively in Julia (I work in finance, we were one of the main sponsors of JuliaCon, though these are my personal thoughts) I want to chime in a little.
To get to 1.0 many of the cons in the language will be solved and there are plans in place for much of this.
I never cease to be surprised by the brevity of code to solve complex problems. Something which would be hundreds of lines in Java/C++ often turns into 50 or so lines of much more readable code.
When asked to describe Julia I struggle but end up this way 'I can make Julia 'dance' like no other fast programming language' though to be fair modern JS is pretty close for many tasks.
The parametric type system is to me the strongest part of the system - I describe it simply as 'templates that work', it is ridiculously powerful and (to me) is the most compelling reason to learn Julia.
I personally find even at 0.4 it is very usable even for tasks which are not math. Though the run time is currently fairly heavy.
So congrats to the Julia team.
Can you elaborate on this?
Re type system, how does it compare to python in use?
The cons, top three from the top of my head, implicit array concat ( going in 0.5 ), String types ( going soon ), anonymous function performance ( gone by 1.0 )
None of these are show stoppers, but they do sometimes fall into the category of surprising. The core dev are very aware of the issues and if you look at the road map for 1.0 many are either well on their way to being fixed or are scheduled to be resolved.
* Non-numeric things are not well optimized. Building a dictionary full of strings was slower than Python when I tried it.
* The module system is confusing because of the principle that files should be "mostly unrelated" to modules. The documentation gives you no advice on how you should organize your code so that it's easy to import, understand, and maintain. The advice I got at a meetup was "just include() everything".
* Package management is a loose wrapper around git. The git commands it runs do not always succeed.
* Curly braces do something different in every version.
Introduction to Julia - Part 1 | SciPy 2014 | David Sanders - https://www.youtube.com/watch?v=vWkgEddb4-A
Get To Know JuliaLang - Guest Talk by Co-Inventor Viral Shah - https://www.youtube.com/watch?v=OC3hsct63Ok
If anyone can recommend other resources that would be appreciated.
The only issue is that, when it comes to numerical computation, there is only "one right way to do it", which goes against Python philosophy. Meaning, either you do it while using vectorial operations mostly everywhere, or it's just too slow to work.
The only real advantage I see in Julia, is that it doesn't constrain you to a specific way to think about calculations. You can think about them in a vectorial way (which is not always possible) or you can just use an iterative approach. Both work fine.
This is what I mainly expect to get from Julia, but we must admit that the inertia behind python and it's excellent libraries for numerical computation, make it hard to change to Julia.
But ... This is actually one of the pillars of Python:
There should be one-- and preferably only one --obvious way to do it
That's like mixing up Vulcans and Romulans. Very dangerous indeed.
There should be one-- and preferably only one --obvious way to do it.[0]
It's strange I know. I work in Python almost every day and I've never touched Perl code.
At least, I guess I'll have a silver lining at cracking my head in vectorizing everything from now on.
I've heard of NumPy, PyPy, SciPy, Pandas, matplotlib, and now Numba. I don't particularly know what these do or how they overlap or interact. Which is kind of the point: to a complete outsider, the world of Python scientific computing feels like a wild west where everybody is happily proclaiming that their setup is just right.
Additionally, Python is slow; Julia is fast. I've heard things like "well, Python is only slow at certain things, and makes it easy to write C code when needed." I don't want to write any C code. Same goes for "typically only a small portion of your Python code is a bottleneck, and it's easy to port that to C."
A huge part of the appeal of Julia is that you don't have to worry about language interoperability, calling C, etc. Everything can be written efficiently and readably in Julia, and because the base language is designed around scientific computing, there's little worry about add-on scientific computing packages not playing nice together.
Now, I fully believe that I could get the right Python environment and set of packages set up be productive. But to get going with Julia, I just download Julia and go.
Also, JuMP [1] is just amazing.
It's not terribly complicated once you spend a little time working with the libraries. Working with arrays? import numpy. Machine learning? import sklearn. Plotting? import matplotlib. Need to do some interpolation, integration, or work with some strange orthogonal polynomials, (etc...)? import scipy. Sure, there's some minor overlap between scipy and numpy, but nothing that causes any problems in my experience.
What I love about using python is that, in addition to all the great math and science libraries, you have all the other python tools at your disposal. Working with xml? import xml. Web-scraping? import urllib2 (or whatever people use now), PyQuery. Additionally, there's all the file system business work that's a joy to do in python using os, sys, etc... libraries.
I'm not sure I buy this argument. Except for pypy, these are all complementary packages and work together. If you breakup anything into its constituent parts, you can make it seem complicated if you want to.
Until you need to do something very straightforward like web scraping or interacting with AWS and find out that nobody's released a package for that yet, so you're going to have to reinvent the wheel before you can get to the scientific question that is of actual interest.
Julia looks very interesting as a language, let's just hope it doesn't end up a ghetto like R (lots of awesome statistical tools, a dearth of libraries and tools for everything else).
For web scrapping: https://github.com/porterjamesj/Gumbo.jl For AWS: https://github.com/JuliaParallel/AWS.jl
Definitely, Julia doesn't have as many libraries as Python, but most common needs are covered. And for the rest you can actually call Python through PyCall.jl
IIUC from GitHub page, AWS.jl has most API implemented, just not thoroughly tested. So it's not really mature or robust, but from my experience with packages at the same stage of development it should be pretty usable for a scientist who "just wants to get shit done".
I totally share your concern regarding marketing solely for scientists, though.
Shit I've seen: silent memory corruption, silently dropping the last element of an array if and only if it is the last declared field in the structure, killing the debugger every time you hit the bridge, inaccessibility of critical code because it uses X unsupported language feature, event-loop integration nightmares (IO suppression / null-routing / deadlock), exception incompatibility, unconfigurable signal/interrupt stealing. And that's all without counting the "usual suspects" of documentation, performance, and testing issues.
Maybe the Julia-Python bridge doesn't have any of these problems. But I'm not going to be the one to find out.
Numpy:
a = np.random.random_integers(low = -100, high = 100, size = (100,))
a[np.logical_and(a % 2 == 0, a > 0)
(or a[(a % 2 == 0) * (a > 0)])Logical indexing is used everywhere in numerical computing. Using functions like logical_and() or boolean multiplication or addition is more difficult to follow than Matlab or Julia.
Julia:
a = rand(-100:100, 100)
a[(a % 2 .== 0) & (a .> 0)]
Well, that's beautiful. Element-wise operations are prefixed with "." a = np.random.randint(-100, 100, size=100)
a[(a & 2 == 0) & (a > 0)]
Python does indeed have logical boolean operators.I think this is clearer than your Julia example because:
- The numpy version makes it clear that you are using random integers, not floating point values.
- The keyword argument for "size" makes it clear what that second 100 is for. (The use of the first two numbers, -100 and 100, is pretty clear from context.)
- In numpy, most operations are element-wise by default, because the result would be ambiguous or not useful otherwise. This removes the line noise of the extra "." before operations.
Don't get me wrong, I think Julia is awesome. I just think you've constructed a very poor example for numpy.
That's also completely clear in the Julia version, if you learn a little Julia.
> In numpy, most operations are element-wise by default, because the result would be ambiguous or not useful otherwise.
This is why I think the Julia approach is better. If I write `a == 7`, am I testing whether `a` is 7 or whether any of the elements of `a` are 7?
In every language I've used, the default is for a "rand" function to return random floats between 0 and 1, and given arguments it returns floats between the arguments. I don't think it has to do with learning Julia, it is just that including "integer" in the function name makes it clear the function returns integers.
> This is why I think the Julia approach is better. If I write `a == 7`, am I testing whether `a` is 7 or whether any of the elements of `a` are 7?
I think this is more of a comment about mixing arrays and scalars in a dynamic language. I made my comment assuming you are performing operations on arrays. If you are comparing two arrays, I think the default of element-wise operations makes more sense.
We started off as a C/C++ library but are constantly adding support for new languages. The one benefit of ArrayFire would be the ability to use the same codebase across CPU, CUDA and OpenCL devices.
>its flexibility and high performance (comparable to C)
This is one of those points that is so rare to see in new, high-level languages, and quite welcome in the world of high-iteration numerical computation.
It seems to be really hard to create smooth user experiences for the common use cases of such complex products. I think people have been trying to oust Matlab for 31 years or so.
Disclaimer: I haven't checked Julia for a few years. Requiring hand reloading of files after editing and no native graph support were a few major problems for a workflow back then.
I totally agree that having dynamic dispatch in the first position of the features list, for example, does not help. But, in my opinion, it is ok for a pre-1.0 language. That is what makes exciting seeing some external help to release a stable version.
It can be a very interesting feature for somebody into language design, but the average Joe working in a lab is not going to learn a completely new language and rewrite all his Matlab code because of it. Most probably, he has not idea what it is or why he would want it.
It was an example of how, at this point, it seems like the web site is more directed towards language developers than potential users.
In this regard, multiple dispatch is actually a huge selling point! I've written in Fortran 90 code, C, Python, and enjoyed most of those languages. There is something unique about being able to pull up the source code for almost every operator/function in the language, implemented in that language. Fortran has this, but is much more statically oriented and difficult for quick scripting. Python, well I'm hopeful pypy keeps building momentum. Or perhaps Pyrhon ported to web assembly, in more pure Python.
Apparently Python is having some success there http://blog.mikiobraun.de/2013/11/how-python-became-the-lang...
http://docs.julialang.org/en/release-0.3/manual/parallel-com...
I am still hopeful. I will give another look at it soon.
I do remember early on when this wasn't the case. Especially when comparing it to Matlab or R. I was discouraged by one professor from learning it. But time has proved otherwise. Julia will likely be the same.
To narrow down, what we're looking for in the immediate future, anybody with interest in core language development, compilers, tool chain support, etc. Additionally, we will probably be hiring for our applications development team very soon, so if you're interested in data science, stats, etc. and like working with people on solving their challenging problems, please do reach out as well.
I'm trying to make an architecture decision right now and this would be helpful (and whet my curiosity :p)
It'd still be a fairly large risk to develop something built-to-last on Julia at this point, so not really the best time for most companies to jump in unless something about Julia really solves a problem they are having.
Jonathan Malmaud, Iain Dunning, Randy Zwitch, Mike Innes (OP) and many others write and maintain general purpose/web-related library code. I'm sure they will agree with some of what I said.
It's interesting that people see multiple dispatch as some esoteric computer science-y thing. In reality, it's just a way of organising code and complexity, just like object orientation is – and it has some compelling advantages over OO as well. For me, it's one of the killer features that I really miss in other languages, and it's well worth taking the time to understand it.
It would be great to see people doing more web stuff in Julia, but for the foreseeable future there will be some important caveats. The web libraries (including in Base) just aren't that fleshed out or battle-tested right now, and there's no Google-scale engineering effort making sure that the runtime is reliable over thousands of CPU hours. Whoever dives into that first will have to have a clear sense of the long-term value.
I agree multi dispatch is awesome, and python has some implementations of it as well.
http://matthewrocklin.com/blog/work/2014/02/25/Multiple-Disp...
If all the logic happens server-side in Escher, what's the responsiveness like? Do you have a live example anywhere?
Are there any plans to transpile Julia into Javascript?
Do you know of anything using Escher for public web-facing applications yet?
Since Julia uses LLVM IR, it should be possible to go from LLVM to Javascript using emscripten.
"more natural for general purpose modeling of data than, say, Python's classes"
Can you elaborate? In most cases single dispatch does just fine. Then class = type except much slower
It's far easier to create new Julia types for modeling a new domain than Python classes. Genrally, types are much more succinct to declare than classes. They're also type checked and parameterized which allow you to do some rather nice things with an API design that would require a lot of if/else behavior checking in Python classes.
The separation of implementation and data declarations via multiple dispatch makes a large assortment of problems much easier to solve. For example, it's much easier to extend the built in behavior of default Julia types such as dictionaries by adding your own custom method (often on line of code) which is incredibly useful for short data processing script.
There isn't as much web focus in Julia currently, but once the implementation makes it easier to possibly "pre compile" an executable I think there could be a big boom in usage for web engines.
I'm on a mobile device currently and can't really write out more explicit details. Maybe later I can write out a lab example.
Based on what I've seen so far, I think Julia could eventually break into the Tiobe top 20 languages (it's currently in the top 100). On the other hand, I think this is plenty. As a language currently aimed at (admittedly fairly large) niches that would be an outstanding outcome.
It also makes it fun to work with as a language. It's not just more chicken.
If an environment is able to have native speed, run in the browser and still be accessible to the people doing scripting now, it could have pretty large implications.
Go also fits the bill I suppose, but Julia is a much better designed language.
This could be overcome but it's not really Julia's niche which is interactive use of notebooks by scientists or other people who are comfortable writing code.
Does anyone think that a dedicated package site for analytics libraries is what is needed ?
From the R documentation:
In the simplest form, an R package is a directory containing: a DESCRIPTION file (describing the package), a NAMESPACE file (indicating which functions are available to users), an R/ directory containing R code in .R files, and a man/ directory containing documentation in .Rd files
But a python wheel is considerably more complicated.
Because I suspect until you acquire mindshare among the academics (who distribute their research as R code), this will be difficult to scale.
[citation needed]
>nobody want to change it.
Love the claim here, followed by a link that reveals an ongoing controversy. True, the core developers don't want to change it, so it's not going to change. They chose to ignore the protests of many would-be users of the language.
I thought that Julia was aiming for Python's niche, i.e. good at numerical stuff, but also decent at inter-operating with other languages.
Argh. We were so close to leaving 1-based indexing behind as a historical relic. But now it seems we'll be stuck with it for a while longer. I'm not saying we can do anything about it now. I just want someone to validate my frustration :-)
OK, this might sound even stupider, but I feel betrayed by the fact that otherwise good, technical people would voluntarily make such a major mistake, and then defend it so vigorously in the face of so many complaints.
Programming languages are used primarily by humans, which means they must be designed with human psychology in mind. Today, Julia would have (even) more community support if not for this seemingly inconsequential decision. The fact that it seems to be succeeding anyway is largely in spite of, not because of, this design choice.
By the way, where can I find perf.h in the benchmark suite? perf.c is there but perf.h seems to be missing.
Does this mean there will be more features limited to DataFrames?
For example, Gadlfy is nice in many ways but by not using Dataframes it means I have to do my own colors for each data I plot. Some of the statistics features in GLM seem to rely heavily on DFs as well.
Dataframes could be much more useful more broadly if the interface was more consistent, better documented (e.g. still haven't figured out how to instantiate with half my data).
Basically, null-ability is pretty useful feature in many fields, not just statistics. Will the plans include any items to help generalize nullability?