The main benefit I can think of is that Julia could theoretically do whole-program optimization, but it doesn't and it doesn't model side effects so even if it wanted to it would be quite limited.
Also writing packages in C++ creates a fair amount of friction. Combining packages becomes easier. The ML world in Julia allows a lot more code sharing while in Python PyTorch and TensorFlow are large monoliths. E.g. different ML packages in Julia can use the same library of activation functions. Each ML library is not a completely different island of functionality.
As a result, despite Julia being very FFI-friendly, the Julia community is trying to move towards a "pure Julia" ecosystem. Sometimes this attitude makes sense (Julia-native BLAS), other times not (Julia has no HTTP/2 packages).
There is not problem if a) someone has written that code in the other language already and b) it works correctly.
But it's harder to build, modify and debug these things, generally. Glue code is useful, but it definitely has a cost.
I was also seduced by the concept of "octave, but with fast loops". This is really all that was needed. But I'm a bit disenchanted by the "advanced" additional features of Julia that add (unnecessary, in my view) complexity to the language. All that multiple-dispatch stuff, all those arbitrary types. I would love a compilation option that forced everything to be of the same type (a multi-dimensional array of numbers), and then it could really run like C, including compilation time.
Well...if you want Fortran, you know where to find it? Fortran still compiles and runs fast. Just use Fortran, then? It's a perfectly fine language.
> if it wasn't for the multiple dispatch thing whose only effect is to slow the compilation (when all you have is the same numeric type).
How does multiple dispatch in Julia slow things down when in Julia (unlike, say, in CLOS) it's basically restricted to generating monomorphic functions on demand? If "all you have is the same numeric type" then only one version of a polymorphic function will get generated and that's as fast as generating code, say, for an equivalent monomorphic function in C.
For one Julia needs an humongous runtime. I can see in the future a compiler from a Fortran number crunching routine to a webassembly SIMD-ready kernel. Julia in the browser is still a failure and not for a lack of effort. Fortran is much simpler to do.
I tried to work in Julia and it has a whole lot of hidden annoying things, undocumented nuances, etc...
I find myself wishing for a language between bare-bones Fortran and big-but-buggy Julia but I know it won't happen since the community is too small
To be fair, Julia needs that runtime precisely to generate compiled versions of polymorphic functions at run time, so as not to cause a combinatorial explosion of n-ary functions. Also, for interactivity, which includes adding new functions at run time. These are problems that Fortran isn't even trying to solve, so it's not quite fair to complain that Fortran doesn't require a runtime that solves them. But of course there seems to be a market for a way to generate a static binary for a Julia program under a closed world assumption that wouldn't require embedding a complete JIT compiler into the process.
That giant runtime enables all sorts of features that are used by unparalleled julia packages.
Soon TM you'll be able to emit runtime free static binaries if you don't need the dynamism. In fact, that's what the GPU compiler does already
Single dispatch is an arbitrary restriction which should not exist. It only exists in most languages as a performance optimization because until Julia came along nobody could do multiple dispatch fast.
It depends on the usage that you make of the language. I don't use Julia as a general-purpose programming language, but only for doing numeric computations. For that usage, the multiple dispatch feature is not necessary (all my variables are matrices of the same numeric type), and indeed it has downsides. For example, the infamous "time to first plot" is so large in Julia precisely due to the need to support multiple dispatch. If this benchmark is becoming faster is only due to a huge and complex effort in optimization. I would prefer if this effort could be spent in improving support for sparse matrices, for example.
I acknowledge that "time to first plot" may be a useless benchmark for many, even for most, people. But you can also acknowledge that multiple dispatch is a useless feature for a few graybeards like me, for whom multiple dispatch is the main reason why Julia feels slow. Indeed, when you use Julia as an interpreted language (e.g., write a Julia script), execution is very slow due to the complexities introduced by multiple dispatch, that force essentially to recompile all libraries upon each execution of my tiny script.
Not to second-guess you (you know far more about Julia than I) but my experience is that people who think that their code is completely uniformly typed still win massively from Julia.
It starts with "just dense matrices with Float64" and then, like the poster above, "but with some nice sparse matrix support" and then "oh and special handling for Vandermonde matrices", "oh and tridiagonals" and eventually "my matrices are hyper-<mumble>-<mumble>-symmetric and I can compute vector dot products really fast". At that point, they have been using multiple dispatch for months without knowing it.
In my own case it was "I don't need multiple dispatch" until it was "oh... it surely is nice that the H3 geo-hashes work transparently with LibGEOS even though the underlying C libraries don't work together."
The point is that Julia makes using things work together far better than anything else I have seen. It even makes completely asocial C libraries talk to each other.
In octave/matlab I can already use the "sum" function over dense and sparse matrices, but you wouldn't say that the language is multiple dispatch. It doesn't seem like a big deal. More importantly, I want the "sum" function to have exactly the same meaning regardless of the data type.
I abhor the idea of hiding algorithms into types, on a very fundamental level. Allowing an algorithm to behave differently depending on the type of the input data seems totally wrong to me. The famous Stefan Karpinski talk at JuliaCon 2019 is one of the most horrifying videos I've ever seen. I still wake up at night trembling and with cold sweat when I dream about this talk.
You really don't. If you have a sparse matrix, you don't want to spend 99% of your time adding zeros.
In general, the advantage of multiple dispatch is it means you can automatically get optimal algorithms for a variety of type combinations. To show why this matters, look at matrix multiplication. If you have a high level function that multiplies matrices, you want to call the appropriate BLAS function (for dense inputs). That function will be one of SGEMM,SSYMM,STRMM,DGEMM,DSYMM,DTRMM,CGEMM,CSYMM,CTRMM,ZGEMM,ZSYMM, or ZTRMM depending on the type (and element type) of matrix (this is a simplified example, in the real world you also might want to diagonal, banded, CSR, CSC, block, or any of 20 or so different matrix types). Without multiple dispatch, you have to either write a bunch of if-else statements to choose the appropriate one, or you just convert everything to a dense (and probably double precision) matrix first. The first one is totally un-maintainable, and the second one will make your program an order of magnitude slower when you ignore structure inherent to your problem. With multiple dispatch, you just call * and it does all the hard stuff for you.
The proof that multiple dispatch is necessary is that most numerics libraries that aren't in Julia make ad-hoc and slow implementations of it internally. For example, here is Pytorch's implimentation https://pytorch.org/tutorials/advanced/dispatcher.html
In case of octave/matlab, of course computing the sum() of sparse matrices does not traverse all the zeros. But the result is the same as if it did, just much faster. There are also many types of sparse matrices, I wonder how does it work internally, it must have a similar mechanism. Probably, since it is interpreted in real time, each time that the "sum" function is called, it checks the type of the arguments and it calls the appropriate function.
> most numerics libraries that aren't in Julia make ad-hoc and slow implementations
The word "slow" has a relative meaning here... if you take into account the time of the first compilation. For example, calling an octave script that performs a matrix product is faster than the equivalent julia script.
I don't know the specifics of how Matlab/Octave do this, but you're probably right about runtime checks.
It's true that compilation time will make Julia slower than Octave for a single multiplication, but if you are doing a bunch of them (especially if they are small), Julia can perform the computations with much lower overhead since it doesn't have to do type checks for stable programs. Also, Julia lets you define specialized matrix types which can be asymptotically faster than Octave for specific circumstances. A great example of this is BandedBlockBandedMatrix from https://github.com/JuliaMatrices/BlockBandedMatrices.jl which are extremely useful solving of PDEs quickly, but aren't implemented in any of the other languages because they have to write all their dispatch systems manually for each algorithm.
Is that correct/can you give a reference? Reading https://en.wikipedia.org/wiki/Multiple_dispatch#Use_in_pract..., julia uses it more in its standard library, but I don’t find indication that Julia is doing multiple dispatch better than, say, CLOS (https://en.wikipedia.org/wiki/Common_Lisp_Object_System) or Cecil (https://en.wikipedia.org/wiki/Cecil_(programming_language)).
Also, what types of dispatch does Jukai disallow that others allowed?
Just curious.
https://arxiv.org/pdf/1411.1607.pdf is a pretty good general resource for info about Julia's design decisions and has a section on multiple dispatch.
The reason that generating more code can be faster is that more code can actually result in fewer cache misses. Julia is so fast because it is really good at inlining, constant prop, and automatic vectorization. These optimizations all increase the amount of generated code, but remove conditionals that decrease code locality.
More code isn't randomly jumped around in, so there are not necessarily (or in practice) more cache misses. Hot paths in Julia compile down better than in many other languages (including C/C++/FORTRAN, where aliasing is much more of a problem), specifically since multiple dispatch gets resolved as much as possible (and often completely) at compile time.
If anything, by making per type versions of things, there are less cache misses since things that are similar stay together, needing less instructions on hot paths to look up types and do other stuff that other systems do hit.
The proof is in the pudding - write (or find) some good code doing similarly complex things, and test them. Julia has C/FORTRAN speeds (and better) for a lot of important tasks, with the development flexibility of Python.
I think one of the nice thing about Julia's "just ahead of time" monomorphization/devirtualization is that it allows a level of dynamism that also works on GPUs/TPUs. This post and linked paper help me understand a little bit of it: https://discourse.julialang.org/t/julia-inference-lattice-vs...
Is this level of dynamism required for conventional ML? Probably not. For physics-informed ML and probabilistic languages? Probably more likely.