Is Julia Really Fast?
medium.com
medium.com
How fast is fast? Well in my case I'm trying to come up with things related to collision detection and I'm getting <30ns using SAT between OBBs (Oriented Bounding Boxes) and <300ns for minimum distance computation using the GJK algorithm again between OBBs.
In the context of recommender systems I did an experiment when I started learning Julia comparing against Cython and wrote about it, although I was a newbie and I should update it: https://plopezadeva.com/julia-first-impressions.html
"Unless you are a C/C++ wizard normal Python developers cannot fix or optimize big libraries like NumPy or TensorFlow. In Julia its users tend to contribute more to the libraries they use themselves."
I am excited by Julia. It seems to have a good niche and community around it. It seems well suited for its job. That focus will probably make it better than more general alternatives in short order.
Easy/Fast C interop is a blessing and a curse. Yes, you can do fast things from C with little overhead, but it also makes things like having a generational garbage collector or a JIT really hard.
Classes in Python are basically sparkling Dicts and their memory layout can change at any time so basically everything must be boxed. Anytime you write a simple loop, the runtime has to ask “okay, what am I looping over!” “what sort of object is this function inside the loop acting on?” “How big is this object?” “Where does it live in memory?”, etc.
This is deeply baked into the semantics of the language, which is why only specific subsets of the Python language can actually be made blazing fast while retaining compatibility with Python semantics.
My understanding was that most JavaScript JITs solved this problem well using shapes.
But it's more than that. There's a really nice talk here from the creator of the Flask framework about a bunch of different semantic choices in the language design, leak out into the ecosystem and make many optimizations impossible: https://www.youtube.com/watch?v=qCGofLIzX6g
What Javascript JITs realized is that even though objects are malleable, they aren't often changed in most code. As a result, JITs make optimizations that assume that the object shape does not change (with checks to deoptimize when that assumption is violated).
But like you point out, this is all thrown out the window because the python C interface effectively directly exposes the memory and structure of python objects to C extensions. The nice thing about that is calling C is super fast and C has a lot of power. It can create, destroy, and modify python objects at will with little overhead.
This becomes problematic with garbage collection because you can't move memory around if C contains a direct reference to that memory (which is where most GC performance benefits come from). It effectively forces python to do reference counting.
Java also has C interfaces but they are MUCH more punishing to use and come with huge performance downsides. Even though they are working to decrease those costs, they won't the lower cost of python C costs. That's because giving a direct reference to an object managed by the JVMs GC is impossible. There is some level of memory copying that simply has to happen.
PyPy is what you get when you relax some of the compatibility constraints of python. It gets pretty significant performance boosts over the default python, but sacrifices the C interface.
Looking at https://github.com/arturofburgos/Assessment-of-Programming-L... https://www.matecdev.com/posts/numpy-julia-fortran.html https://github.com/zyth0s/bench_density_gradient_wfn https://github.com/PIK-ICoNe/NetworkDynamicsBenchmarks
people do find Julia to be faster than Python/Numpy, but it is not uniformly faster than Fortran. And Julia's start-up time should not be ignored. Quoting the last link, "In fact the whole Fortran benchmark (300 integrations) finishes roughly in the time it takes to startup a Julia session and import all required libraries (Julia 1.5.1)."
Using Julia in this case basically means having to rewrite all of the other programs into it and getting rid of eg snakemake or gnu parallel or any number of other very common HPC workflows.
In fact I'd venture that doing a sequence of things many times is just as common an HPC workload than single jobs running for very long times.
This is not to say that the boot time is always a problem, there are plenty of applications in which it doesn't matter at all (as you point out). However that's not always the case (by a long shot) which is why tools like GNU Parallel and snakemake exist.
Julia itself is not even reproducible: https://github.com/JuliaLang/julia/issues/25900 https://github.com/JuliaLang/julia/issues/34753
Not at all. The issue is that the default sysimage distributed with Julia cannot be reproduced across machines, e.g. there is no proof that it hasn't been tampered with. This is an open issue in Guix and Nix, though it affects all platforms.
> One of the places Julia really excels with reproducibility is inter OS
There really is no such thing as reproducibility in Windows, at least not in the sense outlined in bootstrappable.org . If you care about reproducibility, you've chosen the wrong platform.
Different tools for different uses cases is the best way to put it.
Startup time is much improved in recent Julia versions¹, but is certainly not negligible for short calculations.
https://fortran-lang.discourse.group/t/simple-summation-8x-s...
https://discourse.julialang.org/t/i-just-decided-to-migrate-...
Real world systems today are going to use a lot of computing resources, such as clusters, GPUs, tensor processing units, multiple cores etc. In such a world, anything that makes that easy to deal with is going to have the performance edge in practice.
Doesn't matter how fast a Fortran program would be in theory, if the Julia program is delivered years ahead of it.
The AMD [2] and Intel [3] support is younger, but developing quickly.
There's also [4] for a unified API that works across different GPU vendors to avoid lockin
[1] https://cuda.juliagpu.org/dev/
[2] https://github.com/JuliaGPU/AMDGPU.jl
Unless you do a lot of meta programming stuff, I don't find Racket to be advantageous to use. Julia syntax is much friendlier towards general-purpose programming. It is easier to do mathematics as well as string manipulation with Julia.
Multiple dispatch also makes it a lot easier to write code. Getting into the LISP heavy thinking of Racket is harder than jumping onboard Julia.
I find that I have been able to learn functional programming, meta programming and other things that make LISP/Scheme famous in Julia than I ever managed in Racket.
Here's an example: https://ronanarraes.com/tutorials/julia/my-julia-workflow-re...
The naive Julia version has unnecessary allocations and therefore is 23% slower than the optimized version:
@inbounds for k = 2:60000
Pp .= Fk_1 * Pu * Fk_1' .+ Q
K .= Pp * Hk' * pinv(R .+ Hk * Pp * Hk')
aux1 .= I18 .- K * Hk
Pu .= aux1 * Pp * aux1' .+ K * R * K'
result[k] = tr(Pu)
end
In order for this loop to match the C++ version you need to use C-style functions: for k = 2:60000
# Pp = Fk_1 * Pu * Fk_1' + Q
mul!(aux2, mul!(aux1, Fk_1, Pu), Fk_1')
@. Pp = aux2 + Q
# K = Pp * Hk' * pinv(R + Hk * Pp * Hk')
mul!(aux4, Hk, mul!(aux3, Pp, Hk'))
mul!(K, aux3, pinv(R + aux4))
# Pu = (I - K * Hk) * Pp * (I - K * Hk)' + K * R * K'
mul!(aux1, K, Hk)
@. aux2 = I18 - aux1
mul!(aux6, mul!(aux5, aux2, Pp), aux2')
mul!(aux5, mul!(aux3, K, R), K')
@. Pu = aux6 + aux5
result[k] = tr(Pu)
end
... which is quite dirty. But you can write the same thing in C++ like this (and even be a bit faster than Julia!): for(int k = 2; k <= 60000; k++) {
Pp = Fk_1*Pu*Fk_1.transpose() + Q;
aux1 = R + Hk*Pp*Hk.transpose();
pinv = aux1.completeOrthogonalDecomposition().pseudoInverse();
K = Pp*Hk.transpose()*pinv;
aux2 = I18 - K*Hk;
Pu = aux2*Pp*aux2.transpose() + K*R*K.transpose();
result[k-1] = Pu.trace();
}
which is much more readable than Julia's optimized version.If Julia had a linear-algebra-aware optimizing compiler (without the sheer madness of C++ template meta-programming that Eigen uses), then Julia's standing in HPC would be much, much better. I admit that it's a hard goal to achieve, since I haven't seen any language trying this (the closest I've seen is LLVM's matrix intrinsics (https://clang.llvm.org/docs/MatrixTypes.html), but it's only a proposal)
Edit : OK, I see those are small matrices. Then Staticarrays should be a nice contender here, both for speed and readability.
But still, wouldn't even StaticArrays.jl create some unnecessary allocations and copies when written in the naive (clean) way? (for example, when calculating something like z = Ax + By, Ax and By are temporary allocated in Julia's case but you can avoid this in Eigen?)
Even if you used MArrays, which are small statically sized mutable arrays, the intermediate temporary arrays would live on the stack and thus not cause any allocations.
However, there is work happening for large stack allocated arrays in julia. StrideArrays.jl can create stack allocated arrays that work on any size that will physically fit on the stack.
https://chriselrod.github.io/StrideArrays.jl/dev/stack_alloc...
This package is work in progress though and has some rough edges.