Julia: Dynamism and Performance Reconciled by Design (2018)
dl.acm.org
dl.acm.org
My main project is a backtesting framework for algorithmic stock trading. A typical run took more than a month of runtime on my computer when the code was in optimized Numpy and now takes less than 1 minute in optimized (but likely far from optimal) Julia code. The only way to come at all close to the speed of Julia with Python would probably be something like Numba, which I figured would be more immature than Julia. Python is only usable for heavy computing on large data when the glue can be in the outermost loop so to speak.
Numpy / torch shines when the slowest part of the data processing is big matrix multiplication. In that case actually waiting for Julia compilation can slow down experimental iterations.
PyTorch uses tensor cores, supports mixed precision multiplication and bfloat16 representation.
Julia's implementation in CUDA.jl doesn't even use shared memory, which is important, as using that 10MB cache can reduce the needed main memory bandwidth significantly.
Also the new 30xx NVIDIA series supports sparse matrix multiplication, which is great if you have ReLU operations in your data processing for example. I'm not sure if PyTorch has support for that, but I'm sure it will be added, just like how mixed precision multiplication was added this year.
We're talking about an order of magnitude difference if you add all these small things together that PyTorch / cuBLAS supports. Of course Julia can call into these libraries, but then all the nice optimizations that make Julia Julia go away.
I had to think really hard of these things actually, because I love the Julia language, but the more I see ML pipelines beating the previous state-of-the-art models by 5x, the more I understand why an interpreted language with a great linear algebra library won the data processing race (I wish it would have been Ruby though).
Julia also is on the cusp of support for tensor cores, mixed precision multiplication and bfloat16 codegen. New PR just merged last week
Here's an OpenBLAS vs CuBLAS performance comparision. Basically CPU is at this point outdated technology for matrix multiplication. Your second comment is more interesting, I'd be happy if Julia gets to the point when it can beat PyTorch or even JAX on training Lambda networks or Performers.
It's extremely impressive that a high level (using index notation for multiple backends) Julia library can compete with hand tuned kernels. Also bodes well for Julia compiler tech generally that can easily be extended
Julia is going to get to the point of beating pytorch, just taking a but more time due to smaller team and approaching it from a more general position..so when it does it will be more flexible and ergonomic and easily extensible to new techniques, in pure Julia.
1.6 will be a big step as much of the compiler hacks underlying the current ad/gpu codegen (which just fell out accidentally of lispy design) will be replaced with proper tooling for composable compiler passes on typed IR. This will be a phase change imo.
There's already a new faster AD that's almost ready for debut based on 1.6 tech
I stopped using Julia because it was very frustrating that I payed thousands of dollars for a laptop with GeForce RTX 2070 card and the I couldn't make use of it, it felt like I wasted a lot of money...that was my main decision for moving to PyTorch. I know that Julia as a language should be able to beat it easily, but it felt to me that the community was prioritizing CPU over GPU for a long time. I'm happy that it's changing now.
> Basically CPU is at this point outdated technology for matrix multiplication.
This is an incredibly naive statement. There are many many circumstances where you need fast matmul on the CPU even if you have a GPU available because it takes too long to send the memory to the GPU and fetch it once the kernel runs.
If you have many chained matmuls, then yes the GPU is your friend.
That said, regardless of how useful the CPU is in practice for matrix multiplication, BLAS kernels are some of the most overengineered, most optimized pieces of code in wide use.
Tullio is incredibly flexibly and general. The fact that matrix multiplication falls out as a special case in Tullio and can outperform OpenBLAS and match MKL is absolutely stunning.
Before I was working with time series, and there the CPU really shined, but right now I'm classifying sick patients using DNA methylation data (about 100-300 GB / dataset). I'm using a convolutional network that already gives me state of the art results, but I want to experiment with self attention based neural networks. While I'm doing this, most of the biologists are still using logistic regression that can be computed on the CPU.
I see more and more domains (even simulations) where neural networks can outperform classical methods if you know how to apply them, and CPUs can never match the GPU performance. Also as I wrote, the slowest part of all these networks is the matrix multiplication (self attention is especially depending on it).
Pretty much anything where the matmul is important to performance, but doesn't dominate it and requires serial steps. A classic example off the top of my head would be solving a matrix differential equation[1].
Here's an example in Julia using DifferentialEquations.jl and CUDA.jl for the GPU part:
This is solving the differential equation du/dt = A * u for the cases where u and A are arrays on the GPU versus when they are arrays on the CPU. I'm doing the solving in-place to try and minimize the amount of expensive allocations.
using DifferentialEquations, CUDA
function mysolve(u0, A, tspan)
f!(du, u, A, t) = mul!(du, A, u)
prob = ODEProblem(f!, u0, tspan, A)
solve(prob)
end
mysolve (generic function with 1 method)
julia> let n = 50, u0 = rand(Float32, n), A = randn(Float32, n, n)
# Make GPU versions of u0 and A
cuu0 = cu(u0)
cuA = cu(A)
# Do a run of the code so there's no compiler latency biasing the results
mysolve( u0, A, (0.0, 1.0))
mysolve(cuu0, cuA, (0.0, 1.0))
# Now run the compiler code and time it's execution:
@time mysolve( u0, A, (0.0, 1.0))
@time mysolve(cuu0, cuA, (0.0, 1.0))
end;
0.000314 seconds (621 allocations: 98.281 KiB)
0.007218 seconds (10.33 k allocations: 382.266 KiB)
In this case, the CPU code was ~20x faster than the GPU code. True, I could get a better GPU (I'm using a RTX 2060), but I could also get a better CPU (Ryzen 5 2600).There's not much point in using a GPU for this example, and I think that just comes down to the fact that there's a non-trivial amount of serial work that needs to be done between the matmuls.
[1] https://en.wikipedia.org/wiki/Matrix_differential_equation
Of course CPU is faster for small matrices, but again, if I have hundreds of gigabytes of data that you want to process, CPU is always slower.
That is way more data than I am currently using! I'm currently a hobo using data from Yahoo finance. :) That sounds interesting! Could you contact me at kruxigt at gmail.com and tell me more?
The lack of formal rules governing which method will be selected for a given set of arguments has made me somewhat reluctant to use Julia for large projects. Since everything in the language works off of multiple dispatch, it's a little disconcerting to have it governed by vague rules.
However, the manual [0] says that when two methods of equal specificity are present, "Julia raises a MethodError rather than arbitrarily picking a method". The manual doesn't discuss the 5 special cases in the linked PDF, though, as far as I can tell. So did the dispatch / specificity rules get formalized at some point? Or is concern about having formal rules governing dispatch unwarranted?
Of course we do call things like C or Python whenever there’s a good implementation of something we want, and because of that we have some excellent FFI capabilities.
But yeah, all that said, I’d use Julia even if it was slow. It’s an amazing, powerful and fun language that has taught me a lot and made me enjoy programming.
It’s unbelievable that the Julia core devs still haven’t addressed it. It is not fun by any means. Not even tolerable.
> It’s unbelievable that the Julia core devs still haven’t addressed it. It is not fun by any means. Not even tolerable.
While I never really thought this was a big deal, the Julia devs certainly do, and I'm a little upset on their behalf that you'd say this. An amazing amount of developer time and effort has been going into reducing loading times and compiler latency lately. The currently release, 1.5.2 is leagues faster at package loading than the previous version and the upcoming 1.6 release candidate is even faster.
They recently implemented multi-threaded precomilation and have been going on a method invalidation crusade doing largely thankless work to deal with exactly this sort of whining.
Just being honest and voicing my opinions. If it hurts the people that worked on it for years, I am sorry.
-- It's not an opinion when you e.g. say that precompilation has not been addressed when it is being worked on (with much cost, I'd suppose). Same for haven't focused on the UX [snip]. In case you aren't on the julialang discourse I really encourage you to register and ask.
https://docs.julialang.org/en/v1/manual/modules/#Module-init...
Julia 1.6 involved a lot of work reducing invalidations, where old methods are invalidated by new definitions and thus forced to recompile. This has the obvious benefit of reducing the total amount of compilation needed, but the secondary benefit of making compiling to binaries more profitable; it's less likely that work will be wasted when someone loads another package.
https://julialang.github.io/PackageCompiler.jl/dev/examples/...
If it comes to pushing binaries into the package manager, that's what it comes to, but waiting for a minute for the graphing library to load just isn't OK in this day and age.
See this recent blog post for instance: https://julialang.org/blog/2020/08/invalidations/
Ok, that's a little mean, there's nothing quite as convenient and simple as a REPL or notebook for ad-hoc caching, but sometimes you need the opposite and the language really needs to support that.
I can't wait for Julia to murder Python but before it can do that it needs to push all delays out of the interactive parts (to include both REPL and cold launch) and into the package management parts, where people will tolerate delays. Anaconda decided to stick a full SAT solver into the package management and it's every bit the shitshow that you'd expect, frequently taking 30 minutes to solve and failing to converge, but people tolerate it because it's in the package management.
I'm giving it another go though, since it seems like it could be really useful for algorithm-adjacent research.
Closed your terminal session? You're outta luck. Wait 2 mins please. Launch a notebook and want to start working on something? Wait 40 seconds. Wanna check something quick in Julia REPL session? Wait 7 seconds. Let me run this Julia script, wait 55 seconds.
I feel like the entire community should stop doing what they're doing and fix this issue. I am afraid, it is difficult to fix otherwise it would have been fixed by now.
The problem is perhaps that I largely do fairly simple statistics and plotting, with scripts that in R usually take seconds rather than minutes to run, and the disadvantages of the requirement to compile the code overcomes the compiled speed advantage.
https://julialang.github.io/PackageCompiler.jl/dev/examples/...
I’ve never had to write a crazy algorithm. Nor have I had to do any graph traversal. Or anything that impacts performance.
For data analysis, pandas and numpy are plenty powerful. At no time in my career have I said “ditch all this because it’s too slow”. Perhaps others have, but I’m perfectly fine with Python. It’s a beautiful language with ability to hack into basically any part of the language. Batteries, flint, tent and a campsite included.
Python is a pretty poor glue language, in my experience.
Calling C APIs from Python is not pleasant. This is how one might import the libm "error function" from Julia:
const libm = "/lib64/libm.so.6"
erf(x) = @ccall libm.erf(x::Float64)::Float64
Python requires writing/compiling a separate dll written in C (see [^1]). Python extensions are (reasonably) fast, sure. But they are not pleasant at all to write.No, it doesn't, see [0] and [1].
I don't find python beautiful; it's a functional glue language and I don't hate it. It has an ok REPL and some quirky design choices. It has a decent C interop which makes it practical. What makes it useful for data analysis is the large number of packages and staggering number of person-years that have been put into them. The greater python community is great. The language is ok. Together that's pretty useful.
I don't see Julia as a general purpose langauge. I don't see high performance network stack written in Julia ever.
I see Julia as an academic scripting language that happens to be fast *
* Debatable.
For example it might be possible to AoT compile parts of the code (as long as the types are fully defined) to export as a standalone library, and if that happens it would be beneficial to write the number crunching libraries for glue languages in Julia. I mean, there are already those (like DiffEq to Python and R), but if it's a binary that could have just as easily come from C it would be even better. And there are already works along this path, including for wasm deployment.
The multithreading in Julia is very fast and easy to use, and I'm sure it will become robust and safer as well as the language gets older, which can very well allow for an Akka style framework that allow for large software with both high reliability (thanks to stuff like supervision trees) and structured code (since like Julia's ecosystem design favors small libraries composing with each other, the future of large Julia codebases could become the composition of many small high level but fast actors).
There is also lots of interests in work-arounds on the garbage collector, like improved support to immutable structures (like 1.5 improvements on structs with references to mutable structures). It's possible that you might be able to more reliably write parts of the code that do not allocate (or allocate in predictable ways), so you'll be even handle those realtime code loops.
Sure this is very speculative, but look at Python today compared to when it had Julia's age. It took 8 years since Julia became public to get to this point (2 years since it got the first stable release, 1 year since it got it's multithreading model), it will surely be different when it has triple that age.
At present, the Python ecosystem -- packages, frameworks, open-source repositories, tutorials, online discussions, communities, documentations, etc. -- is a significant advantage in favor of Python.
Fully agree, but with PyCall you can just use any Python package in Julia: https://github.com/JuliaPy/PyCall.jl
I do that regularly when I need something from Python and it works perfectly. So, Julias ecosystem is really a superset of Pythons when it comes to libraries.
Standard ML will catch up eventually
I am not sure if it supports plotting libraries à la matplotlib, however. If it has this capability, it might be still interesting in its present state.
[0] https://www.youtube.com/watch?v=Qa_9vut4TzQ&list=PLxLdEZg8DR...
There are a few tests which have Julia and Nim implementations. It would be nice if the benchmarks game maintainer would include some more languages, like Nim, Crystal, and Zig.
This doesn’t have Nim, but is another interesting one. https://modelingguru.nasa.gov/docs/DOC-2783
(Name-calling isn't a good look.)
If they did nothing to be findable they probably wouldn't be.
I dare say that's less effort than trying to make something be successful.
Constantly doing the work is plainly not zero cost.
Python with Numpy and JIT, as well as Matlab, R and FORTRAN could also be included with suitable tasks. The code should be optimized in each case.
https://julialang.org/benchmarks/
AFAIK these have been optimized for the other languages, but your opinion may differ if you look at the implementations. Code for these is available here:
These things are all pretty subjective though when it comes to 'what is a fair benchmarking methodology'? Another valid approach would be to have experts in each language try to write as optimal an implementation as possible, but this would quickly become a bit of an arms race.
Come on!
https://github.com/dyu/ffi-overhead
This is one reason why you could see Julia outperforming Fortran in some cases where the FFI speed matters. But Fortran does have easier aliasing analysis (because you can't alias) so that helps there, but other than that most of the compiler passes are pretty much the same.