Julia: Performance Tips
docs.julialang.org
docs.julialang.org
Looking back, I think the tooling Julia had from the start, combined with the REPL, made it actually really nice to fix these performance issues. Much better than compiling and linking a binary and running it through some heap profiler tool. You could simply do
julia> @time f()
x.y seconds (N allocations: M bytes)
and iteratively improve part of the code base instead of profiling the entire application.(To be fair: back then escape analysis was not implemented in the compiler, and it was hard to avoid silly allocations)
Note: I don't know Julia specifically, but this does apply to other GC languages like Ocaml and Java. Try reading code in those languages that avoid the GC. It looks very strange. In python it is even worse: people try to avoid for loops because they are slower than list comprehensions for example (or at least used to, I haven't written much python for some years now).
Generally, the code ends up looking rather similar to non-GC languages. You create some buffers outside of your performance-sensitive parts, and then thread them through your code so they can be accessed and re-used in the hot loop or whatever.
It could be better, e.g. C++ and Rust both have some nice utilities for this stuff that are a bit hard to replicate in Julia, but it's not auch a huge difference, and there's also a lot of advantages on the julia side.
E.g. it's really nice to have the GC available for the non-performance critical parts of your code.
Because from what I have seen in other GC languages, the answers to any of those questions haven't been great.
You can profile memory usage line by line in detail pretty easily with tools like @timeit or @btime
In practice I've found it pretty easy to get inner loops down to few or 0 allocations when needed for parallelization
There's static analysis tools that might be over-eager finding allocations, and there's also runtime measurement tools which should go into your test suite if it's of vital importance to monitor.
Turns out a little bit of ergonomics actually matter.
We've ported some tens of thousands of lines of numpy-heavy Python, and in practice our Julia code is actually more concise while being about 10x-100x more performant.
Sure, but why not write it all in Rust or similar then? (Not writing it all in C++ I would understand.)
> This is especially nice in situations where you expect to throw away a lot of the code you write, e.g. research.
Right, that is very different from what I do. There is code I wrote 15 years ago that is still in use. And I expect the same would be true in 15 years from now. Though that is also code where a GC is a no-go entirely (the code is hard real-time).
I don't know Rust as well as I'd like, but when we've worked with some strong Rust programmers, their versions of the code are something like 4x as long as equivalent Julia code with minimal performance improvements. And since we primarily care about trading algorithms that don't have binary results, it's pretty helpful to be able to understand those at a glance.
Also, Rust's ecosystem for numeric computing still seems pretty underdeveloped, though it's getting better.
Because it is clunky, ugly and unreadable mess. Which is something you may be willing to pay in some circumstances, but not while researching algorithms, doing data science, writing simulations etc.
> Though that is also code where a GC is a no-go entirely (the code is hard real-time).
There are GC systems with hard real time guarantees. There is more out there than just OpenJDK and Go.
Also, hard real-time systems rarely use runtime allocations, at least with off the shelf allocators.
D, C#, Swift, Nim,.....
Agree that in Julia's case the flexiblity is not quite there, still much better than using Python and then going to write most of the work in C, C++, Fortran,.....
Which is a thing that gets lost quite often in these discussions, just because the last 5% might be a bit harder, doesn't mean we have to throw everything away and start from scratch in another programming language, with its own set of problems.
Well, a lot of C/Odin/Zig people will point out that Rust's stdlib encourages heap allocations. For actually best performance you typically want to store your data in some data-oriented data model, avoid allocations and so on, which is not exactly against idiomatic Rust, but more than just a typical straighforward Rust just throwing allocations around.
As for data oriented design: yes, and it depends on your usage pattern. E.g. arrays of structs can still be better than struct of arrays if that matches your cache line access pattern. Zig probably makes SOA easier for the cases where that is more efficient (I haven't used Zig, but I have read a bit about it. It has some cool ideas (comptime), but also things I can't get behind such as lack of RAII, and generally lack of memory safety.)
In retrospect, the PhD experience was miserable and using Julia contributed to my 'death by a thousand cuts'.
Some people may not realize it, but when it comes to programming languages, ergonomics matter—a lot.
Why are you recalculating indices?
I don’t think that any Julia program I’ve ever written would need to change if Julia adopted 0-based tomorrow. You don’t typically write C-style loops in Julia; you use array functions and operators, and if you need to iterate you write `for i in array ...`.
“ergonomics matter”
Definitely. Ergonomics is the main reason I enjoy Julia. Performance is a bonus.
[1] https://gerritnowald.wordpress.com/2022/10/03/simulating-rot...
Our Julia code is parallelised with FLoops.jl, but so far Numba has shown surprising performance benefits when executing code in parallel, despite being slower when executed sequentially. Therefore I can imagine that Julia might yield better results when run in a regular desktop environment.
https://github.com/JuliaParallel/rodinia/tree/master/julia_m...
It was touched 9 years ago, but maybe you have ported it to current standards. I don't think we had multithreading at that time, only multiprocessing.
Is your Julia implementations available somewhere? (Sorry if it is in your paper but I missed it). I vaguely remembered in the past that working with threads leaded to some additional allocations (compared to the serial code). Maybe this is also biting us here?
As far as I know the code was ported to use @floops, with minor optimisations in addition to that.
I think it's quite possible that it's an allocation issue, that's something we're looking into, although I don't have any specific results for Julia yet.
To bring Julia performance on par with the compiled languages I had to do a little bit of profiling and tweaking using @views.
I am not familiar with Julia nor Numba internals, but maybe Numba, due to being more specialized, can actually provide LLVM with info that allows it to make more aggressive optimizations more easily.
I have a suspicion that Julia, owing to multiple dispatch, has a sort of regularity that makes that you said plausible.
Though there is just so much more Python to train on, any I bet they even do RL with validated rewards on Python, and probably not Julia.
It is also an excellent language for messing about because the language and especially the REPL have tons of quality-of-life features. I often use it when I want to do something interactively (eg. inspect a data set, draw a graph, or figure out what's going on with an Unicode string, or debug some bitwise trickery).
What Julia is not great at is things where you need minimal overhead. It is performant for serious number crunching like simulations or machine learning tasks, but the runtime is quite heavy for simple scripting and command-line tools (where the JIT doesn't really get a chance to kick in).
(I'll take strong static typing every day, it is so much better.)
Though poorly documented APIs exist everywhere, but they are not something you can rely on anyway: if it isn't documented the behaviour can change without it being a breaking change. It would be irresponsible to (intentionally) depend on undocumented behaviour. Rather you should seek to get whatever it is documented. Otherwise there is a big risk that your code will break some years down the line when someone tries to upgrade a dependency. Most software I deal with is long-lived. There is code I wrote 15 years ago that is still in production and where the code base is still evolving, and I see no reason why that wouldn't be true in another 15 years as well.
At least you should write tests to cover any dependencies on undocumented behaviour. (You do have comprehensive tests right?)
I usually write the tests afterwards, except for very well-defined engineering problems, and the REPL exploration helps inform what tests to write.
To produce plots out of data files, Python and R are probably the best solutions.
Makie has the best API I’ve seen (mostly matlab / matplotlib inspired), the easiest layout engine, the best system for live interactive plots (Observables are amazing), and the best performance for large data and exploration. It’s just a phenomenal visualization library for anything I do. I suggest everyone to give it a try.
Matlab is the only one that comes close, but it has its own pros and cons. I could write about the topic in detail, as I’ve spent a lot of time trying almost everything that exists across the major languages.
Also, the REPL. Julia's REPL is vastly better than any other language REPL, by far. Python's is good but Julia is way better, even as a calculator for instance, it has fractions and it is more suited for maths
Julia won for a very simple reason: we tried porting one of our major pipelines in several languages, and the Julia version was both the fastest to port and the most performant afterwards. Plus, Julia code is very easy to read/write for researchers who don't necessarily have a SWE background, while Go or C++ are not.
We started using Julia in the Research Infrastructure team, but other teams ended up adopting it voluntarily because they saw the performance gains.
So better question is: in which circumstances would you choose Julia over more mainstream-y alternative like Clojure? And here scientific and numerical angle comes to play.
At the same time I think Julia is failed attempt, with unsolvable problems, but it is a different topic.
Yes sure Julia isn't fully dynamically typed, but that doesn't change the fact that it isn't fully static typed. If it was, it should be pretty easy to create static binaries and find bugs like "func not defined" at compilation time.
1. Julia has great tooling for operations research/linear programming. JuMP provides an standardise interface to interact with solvers (e.g., Gurobi, CPLEX) via wrapper libraries.
2. I like its overall ergonomics. It is fast enough that a programmer might not need to use a compiled language for performance. The type system allows for multiple dispatch. And the syntax is more approachable than say Python for matrix algebra.
3. I would say the performance is overstated by the community but out of the box it is good enough to avoid languages like C/C++ to build solutions. The two-language problem in academia is real, and Julia helps to reduce that gap somewhat in certain fields.
A very minor nit: Julia is a compiled language, but it has an unusual model where it compiles functions the first time they're used. This is why highly-optimized Julia can have pretty extreme performance.
> I would say the performance is overstated by the community but out of the box it is good enough to avoid languages like C/C++ to build solutions.
For about a year we had a 2-hour problem in our hiring pipeline where the main goal was to write the fastest code possible to do a task, and the best 2 solutions were in Julia. C++ was a close third, and Rust after that.
Or anyone living in any JIT ecosystem for that matter, when it actually kicks off is an implementation detail, and some ecosystems even have configuration options.
Caught. Should have just listed the usual suspects (C, C++, maybe Rust nowadays?).
> and the best 2 solutions were in Julia. C++ was a close third, and Rust after that.
Awesome. Which type of problem was this, if you can share?
The C++ versions were around ~250 lines long, and the developers generally only had time to try one approach. While the Julia versions were around ~80 lines long, and the developers had time to try several approaches. I'm sure the best theoretically possible C++ version is faster than the best Julia version, but given that we're always working with time constraints, the performance per developer hour in Julia tends to be really good in my experience.
https://docs.julialang.org/en/v1/manual/mathematical-operati...
Go is similar in many ways, but takes a performance hit in areas like Garbage collection.
The Julia community is great, and performance projects may also be compiled into a binary image. =3
https://julialang.github.io/PackageCompiler.jl/dev/devdocs/b...
They greatly improved the generated binary file size last year. ymmv with "--trim=unsafe-warn" =3