New horizons for Julia
lwn.net
lwn.net
the `--trim` feature coming in 1.12 will still be considered experimental, and if I had to take a guess won't be considered "stable" until 1.14 or so, but I'm really excited to see it picked up around the ecosystem!
The syntax is easy for Math-heavy Researchers who might not have as much software engineering experience to use, but it's much faster than numpy, which we used for 5 years before.
The package management is sane – not as good as Rust, but way better than Python/Javascript.
We have one non-trivial macro which makes our codebase ~15% shorter, and would be hard to emulate in a lot of other languages.
The REPL is really good, second only to Clojure IMO.
The main thing I want to see more of is "static Julia", e.g. better support for static type-checking and binaries, and fortunately that seems to be where the language is going!
This! I’ve found a lot of value in Julia for math specially in financial analysis.
Julia’s DataFrame and general development workflow are excellent for research and when it comes to deployment is not too shabby either.
I’m aware of pandas and the rest in Python, but for some reason Julia feels a lot more “ergonomic” for my use case.
It's worth noting that we initially migrated ~15kloc of numpy data/numeric pipelines to Julia, and mostly found the same values, with the exception of something that turned out to be a bug in a Python library.
I would say that the biggest Julia annoyance we've run into has to do with the way the Expr type is implemented, particularly the fact that it's mutable, which makes some of the metaprogramming we want to do substantially harder. But that's not a bug per se, just a design choice I don't like.
* usage of `OffsetArrays.jl` (really, just avoiding this package fixes most of the issues)
* hideously cursed syntax that should only pass a code review if your name is Lovecraft, but for some reason the parser allows it
if you don't do either of those things (which at least personally speaking, I don't) then I don't think the rate of "correctness issues" is any higher or lower than I experience in other ecosystems. In fact, it's probably lower
Don't forget that other languages are not immune... I love the `polars` library as well but in my two years of using it I've encountered organically two separate "correctness" bugs. It's just par for the course for any big code surface
In addition, I believe that abstract types are not that horrible since in industry we use OOP any way (https://github.com/Suzhou-Tongyuan/ObjectOriented.jl). I coworked with the author of this package several years ago. Currently, he is developping a different branch of Julia compiler. Since OOP makes eveything easier to design (from linter to static compiler), and programmers prefer OOP over abstract types, I personally don't think they will cause huge problems.
But my main complaint about Julia is its general approach to memory management. You are encouraged to think about how memory is being allocated e.g. by using StaticArrays for smaller, immutable arrays, or pre-allocating the arrays and operating on them in-place. But you don’t actually have control over allocations, and it’s easy for some type instability to cause unnecessary allocations - even after you stick a bunch of type annotations which should give the compiler sufficient information to force type stability or at least crash/fail to type check.
At this point I think people are better off investing their time/efforts into Rust for computationally heavy workloads on the CPU (polars for data frame stuff, ndarray for numpy functionality, Enzyme for autodiff) and using torch/jax for GPU-centric work (also, it’s significantly easier to write python bindings for rust libraries than C++ or Julia libs).
I'm not a full-blown hater, but I have problems with that as well. Specifically, you have no control about it whatsoever, you're just promised that "if you do things right, it'll be amazing". And it is! The problem is that any tiny minuscule mistake causes catastrophic failure of performance due to allocations. Since the good performance depends on type stability, and type stability propagates, any mistake anywhere will propagate everywhere. Think: if a variable becomes type unstable due to a programmer mistake, any function that consumes it generally might become type unstable as well, and any function that consumes the output of that function as well, etc. The upshot is that this forces you to think more carefully about your types and data structures. Programming in Julia extensively has made me a better programmer. I'm not a C++ expert, but I believe that in C++ these kind of mistakes always end up being localized.
It's true that accidental dynamic behaviors are a real concern and can be a performance killer. Fortunately, the language has nice tooling. In VSCode, I often use the visual profiling tool `@profview` to get a flame graph. Anything dynamic gets highlighted in red, and is quick to diagnose. There also exist nice static analysis packages like JET.jl. During development, one can use `report_opt` to statically rule out accidental dynamical behaviors. Such checks can also be incorporated into a project's unit tests. In practice, it's not much of an issue for me anymore. But to be fair, there is a big learning curve for new Julia users. See, e.g., https://docs.julialang.org/en/v1/manual/performance-tips/
At the same time, I would expect the Rust ecosystem to overtake Julia’s in that domain in the next couple of years. Polars is already nicer than pandas, I’ve seen a some promising work on numpy-style tensor libraries, and I’m pretty impressed by the progress with getting enzyme integrated into Rust (I could never make it work with Julia). Here’s a nice example repo I saw recently:
I think something in between Zig and Rust will emerge someday as a sort of optimal compromise between compile speed, safety, and programmability wrt memory and performance tradeoffs.
Also, it's worth emphasizing that the user experience of Julia has been improving greatly, even in just the last 3 years. Julia 1.9 introduced caching of native code [1], and now at Julia 1.11 the time-to-first-plot in a new Julia process is typically less than a second.
Having said all this, Rust is an absolutely fantastic language too, and might be preferred for large-scale software development efforts where static analysis is prioritized over an interactive development workflow.
[1] https://julialang.org/blog/2023/04/julia-1.9-highlights/#cac...
Why not and how so?
- Julia's lack of formal interface specification. Julia gets a lot of flexibility from its multi-method dispatch. This feature is often considered a major selling point in allowing code reuse. Many Julia packages in the ecosystem can be combined in a way that "just works". Consider, for example, combining CuArrays.jl with KrylovKit.jl to get GPU acceleration of sparse iterative solvers (https://github.com/Jutho/KrylovKit.jl/issues/15). But it's not always clear who actually "owns" such integrations. Because public interfaces aren't always well documented in Julia, things are prone to breakage, and it can sometimes feel like "action at a distance". This was especially painful with the OffsetArrys.jl package, which suddenly introduced arrays that could begin at any integer index. (That was the major theme of Yuri's blog post, and the simple solution for most people was to avoid OffsetArrays.) Rust's community philosophy and formal trait system err on the side of providing static guarantees for correctness. But these constraints also take away flexibility to fit distinct packages together. For example, Julia has always had excellent support for type specialization, and this has been notoriously challenging to fit into Rust, even in a very limited form: https://users.rust-lang.org/t/the-state-of-specialization/11.... Conversely, there have been many discussions about designing a formal interface system in Julia, but it remains a challenge: https://discourse.julialang.org/t/proposal-adding-optional-s...
- Julia is designed around just-in-time compilation. For example, every time a function is called with new argument types, it will be freshly compiled for that specialization. This is great when you care about getting optimal performance. Also, because Julia allows to reify syntax as value-level objects, you can assemble Julia code that is custom optimized to run-time values. All of this is amazingly powerful for certain kinds of number crunching codes. But carrying around a full LLVM system is clearly a blocker for distributing small, precompiled binaries. Hence the LWN discussion about the preview juliac feature, which will offer a mode for fully static compilation.
- Rust's borrow checker is something to envy. In any other language, I miss the ability to safely passing around references to stack allocated variables, or to know that a referenced value cannot be mutated.
Finally, I would probably recommend Python (not Rust!) for most machine learning or data analysis projects that aren't too "bespoke". There's just so much momentum behind PyTorch and JAX. The Julia community is developing some very interesting packages in this space. Notably, Lux.jl, Enzyme.jl, Reactant.jl, and all of SciML. These are super powerful, but still very researchy. For simple things, Python will probably be less friction.
The best language will depend on your use case. Julia serves its niche very well, even if it doesn't fit every possible use case.
There was also some other if I remember correctly I saw it in the comments of https://youtu.be/eRHlFkomZJg, but now I don't see it , I think it was evcxr only.
There is this if we really want a rust repl
While I always had an eye on Julia and it is a lovely language, especially the multiply dispatch and general lispy-ness in spirit, I never really had a good use case for it.
Python always was the better choice due to the stronger ecosystem and faster launch time. Rarely did I need the extra performance Julia offered. But Python is annoying to deploy so Julia starts to seem tempting, at least for smaller projects where you don't want to do a complex docker setup.
Julia is for sure prioritizing the right things.
You can also upload it to pypi for example with name let's say for example heytest and then just do uvx heytest and it would just work.
Uv is great , uv is love.
I think this is about right. If you are going to solve differential equations or run some numerical weather codes or solve an integer program or simulate an electrical circuit, it would be a good idea to consider Julia rather than reaching for Fortran, C++, or MATLAB. It is a much nicer language and still very fast.
My experience is that people who want to use Julia for data science and especially for deep learning, usually find that you can get the job done, but often it's easier to use PyTorch or Jax and just accept the limitations of Python.
However, I have read enough Fortran and C++ written by scientists, to fervently wish that their code had been written in Julia instead. In Julia, I don't think they could manage to write code so debauched. If you just do things the naive way, usually life is good and code is both readable and fast.
The pitfall is type instability and memory allocation, but I think with a few basic tricks, people can learn to avoid this, especially for simpler tasks. (edit: also, if people do mess this up, they notice because their code is slow. whereas people don't notice if their code is fast but incomprehensible and impossible to maintain or extend.)
currently if in a file i have
local s = 0
for i = 1:10
t = s + i
s = t
println("$s")
end
println("$s")
and I execute this file, I will get an error ERROR: LoadError: UndefVarError: s not definedYou can write this as either
s = 0
for i ∈ 1:10
t = s + i
global s = t
println("$s")
end
println("$s")
or let s = 0
for i ∈ 1:10
t = s + i
s = t
println("$s")
end
println("$s")
endSure, there's things where I find it too fussy and I wish it'd go "yeah yeah I get what you mean" and there's also things where I find myself thinking "man, I really wish you had complained to me about that, it was clearly a bug / bad style!"
But all-in-all, I find it's a language that mostly gets out of my way and lets me code, which I find very useful and nice for scripting.
I think if one finds things like scoping rules annoying for scripting, that might just come down to just knowing and being more familiar with another language that has different rules.
Why not make this the default everywhere? Well, there are a lot of scientific use cases where it's convenient to have Python-style dynamic typing and interactivity. A cool thing about Julia is that it allows that, while _also_ allowing to achieve high-performance, all within a single language.
For the record, I do also love the static type system and overall design of Rust. But for my day job (research in numerical methods and computational physics), I find Julia to be the most efficient way to get the job done -- rapid algorithm prototyping, data analysis, plot generation, etc.
function test()
s = 0
for i = 1:10
t = s + i
s = t
println("$s")
end
println("$s")
end
test()
For anyone curious, the rationale for this scoping is explained in the docs. Essentially it's to prevent the so-called spooky action at distance.s is defined in the statement and then is undefined out of that statement because there is no other scope for it - it's at top level. just write s =0 and all is well as s exist in the global scope where your code lives as well...
[1]: https://docs.julialang.org/en/v1/manual/variables-and-scopin...
Although it's a bit of an challenge to get the interpreter to understand which thing that's out of scope should be in scope for the code to work (because it's all out of scope..)
Julia biggest problem at the moment is growth. Julia has suffered from not having exponential growth, and has either maintained a small linear growth or has fallen in popularity. Search online on YouTube for tutorials, on Twitch for WatchPeopleCode, or on GitHub for benchmarks; and Julia is not even in the room where the conversation is happening - there just isn't any mindshare.
And for good reason. There are so many ergonomic challenges when using Julia in a large codebase and in a large team. Julia has no formal interfaces, LSP suggestions that are often just wrong, and no option types. This just makes writing Julia code a drag. And it makes it quite difficult to advocate to developers experienced with languages that offer these features.
Additionally, the core conceit pushed by Julia advocates is that the language is fast. This is true in controlled benchmarks but in real-world scenarios and in practice it is a real pain to write and maintain code that is fast for high velocity teams because it requires a lot of discipline and a strong understanding of memory allocation and assumptions the Julia can and cannot make. You can write code that is blazingly fast, and then you make a change somewhere else in your program and suddenly your code crawls to a halt. We've had test code that goes from taking 10 minutes to run to over 2 hours because of type instability in a single line of code. Finding this was non-trivial. For reference, if this were uncaught our production version would have gone from 8 hours to 4 days.
The lack of growth really hurts the language. Search for pretty much any topic under the sun and you'll find a Python package and possibly even a Rust crate. In Julia you are usually writing one from scratch. Packages are essential to data processing are contributor strained. If you have a somewhat unpopular open source code code you rely on that doesn't work quite work the way you want it to, you might think I'll just submit a PR but it can languish for months to a year.
The Julia community needs to look at what programming languages are offering that Julia developers want and will benefit from. The software world changing very quickly and Julia needs to change too to keep up.
For those who might not be familiar, tooling can sometimes help a lot here. The ProfileView.jl package (or just calling the @profview macro in VSCode) will show an interactive flamegraph. Type instabilities are highlighted in red and memory allocations are highlighted in yellow. This will help to identify the exact line where the type instability or allocation occurs. Also, to prevent regressions, I really like using the JET.jl static analysis package in my unit tests.
I just love pythons ease of use, and the few times I've tried Julia it has felt clanky in comparison, and then Jax just makes everything so amazingly fast.
the language is at its best in other areas. for a few things (differential equations, mathematical optimization) it is actually state of the art.
I don’t think that Julia is a great tool to replace the complicated physics simulation code that would have traditionally been written in C++/Fortran, I think for that you’ll generally be better off using JAX/Rust/C++.
I would be thrilled if Julia had this in early days, but now it is a little too late. I have jumped the boat.