Julia aims to be better than R and Python at statistics and data analysis. It's not there yet, but I could easily see it replacing a great deal of academic use of Numpy and Python in Jupyter notebooks (the 'ju' is Julia).
On the other hand, Rust seems like it's aiming at being a safer alternative to C for low-level systems programming.
I'm sure there's things Julia could learn from Rust, but the design decisions are going to differ wildly because they just aren't trying to do similar things.
Because it's so closely integrated with C, Python actually does a decent job at this, which IMO is one of the reasons it's been so successful as a general purpose language in addition to a scientific one. I really hope Julia can do the same.
Off the top of my head, here are a few features I think would really improve Julia for general purpose programming:
- Static type checking (JET.jl[1] looks promising)
- Faster startup (again, great strides have been made)
- Tagged, closed unions
- Pattern matching
MLStyle.jl [1] is quite nice for this and has been around for a while.
> Tagged, closed unions
These are less general than 'real' unions and can be implemented using them. E.g. SumTypes.jl [2] has some macros to make it a bit more convenient to define them, it could use some other quality of life features though.
[1] https://thautwarm.github.io/MLStyle.jl/latest/syntax/pattern...
For tagged unions, the difference is in memory layout. A real tagged union type would eliminate indirection and allow more code to be type stable.
For instance,
mutable struct NotTypeStable
o::Union{Some{Int}, Nothing}
end
function not_type_stable()
o = returns_option()
if isnothing(o)
0
else
1
end
end
versus: enum Option{T}
Some(T)
None
end
mutable struct TypeStable
o::Option{Int}
end
function type_stable()
o = returns_option()
match o
Some(_) => 1
None => 0
end
end
Both of these features aren't strictly necessary, but they act as a compliment to Julia's existing dynamic mechanisms (multiple dispatch & Union types) julia> @code_warntype not_type_stable()
MethodInstance for not_type_stable()
from not_type_stable() in Main at REPL[2]:1
Arguments
#self#::Core.Const(not_type_stable)
Locals
o::Any
Body::Int64
1 ─ (o = Main.returns_option())
│ %2 = Main.isnothing(o)::Bool
└── goto #3 if not %2
2 ─ return 0
3 ─ return 1
julia> versioninfo()
Julia Version 1.7.0-DEV.1169
Commit e5d7ef01b0* (2021-05-26 14:17 UTC)
Note `Body::Int64`.
With older Julia versions, it should still be type stable if you did `o === nothing` instead of `isnothing(o)`. using MLStyle
mutable struct TypeStable
o::Union{Some{Int}, Nothing}
end
MLStyle.@as_record TypeStable
MLStyle.@as_record Nothing
function type_stable()
o = returns_option()
@match o.o begin
Some(_) => 1
Nothing() => 0
end
end
returns_option() = TypeStable(rand(Bool) ? Some(rand(1:10)) : nothing)
And what does the compiler have to say about it? julia> Core.Compiler.return_type(type_stable, Tuple{})
Int64
and julia> @code_warntype type_stable()
Variables
#self#::Core.Const(type_stable)
o::TypeStable
259::Union{Nothing, Some{Int64}}
return#257::Union{Nothing, Int64}
Body::Int64
1 ─ (o = Main.returns_option())
│ (return#257 = Main.nothing)
│ (259 = Base.getproperty(o, :o))
│ %4 = (259 isa Nothing)::Bool
└── goto #3 if not %4
2 ─ (return#257 = 0)
└── goto #5
3 ─ %8 = (259::Some{Int64} isa Some)::Core.Const(true)
│ %8
│ %10 = (259::Some{Int64} !== Main.nothing)::Core.Const(true)
│ %10
│ (return#257 = 1)
└── goto #5
4 ─ Core.Const(:((error)("matching non-exhaustive, at #= REPL[9]:3 =#")))
5 ┄ return return#257::Int64
This is on version 1.6.1 for me, but should work fine on earlier releases. I agree that having MLStyle.jl pattern matching bundled into julia would be great!No, of you look at the typed IR I posted, it says
259::Union{Nothing, Some{Int64}}
and (259 = Base.getproperty(o, :o))
So it's inferred properly to be a small union which julia handles efficiently at runtime: julia> @btime type_stable()
18.988 ns (1 allocation: 32 bytes)
1
> Furthermore, TypeStable itself has to be on the heap, which I think means this causes two pointer lookups rather than just one.It's on the heap because you wrote
mutable struct
in your example and I was just mimicking you as close as possible. If you don't actually want it to be mutable, then you can just remove the keyword mutable and it'll be even faster and have no heap allocations: julia> @btime type_stable()
18.426 ns (0 allocations: 0 bytes)
1Oh! I didn't realize Julia already optimized small unions.
https://julialang.org/blog/2018/08/union-splitting/ https://docs.julialang.org/en/v1/devdocs/isbitsunionarrays/
I wasn't quite sure whether something like this was possible given the semantics of Julia.
Personally, I'd still prefer something more explicit, so it would be more obvious when a union is being handled efficiently. But given that small unions are optimized, I can see why Julia would make the opposite decision.
This makes me feel a lot better about writing Julia code with union types. Thanks
There's even a numerical computing library: https://ocaml.xyz/
The biggest reason is because some function of the high level language is incompatible with the application domain. Like garbage collection in hot or real-time code or proprietary compilers for processors. Julia does not solve these problems.
Other reasons are practical, like portable executables in Go or Rust (portable in the sense they do not require dependencies on their target systems, usually). Julia does not solve this problem.
Then there are the reasons to use scripting languages over the system languages, like expressive syntax with low cognitive overhead. Julia definitely helps here. But this comes at the cost of execution and startup time. Julia only kind of solves this problem.
So if Julia is trying to make a more performance scripting language then that is admirable. But for most of the projects where I have needed to prototype in a script and implement in a systems language, Julia would not have worked.
Even today I don't have a good reason to use it over MATLAB for day to day work, since I already have the license and their ecosystem is more mature.
The presence of garbage collection in julia is not a problem at all for hot, high performance code. There's nothing stopping you from manually managing your memory in julia.
The easiest way would be to just preallocate your buffers and hold onto them so they don't get collected. Octavian.jl is a BLAS library written in julia that's faster than OpenBLAS and MKL for small matrices and saturates to the same speed for very large matrices [1]. These are some of the hottest loops possible!
For true, hard-real time, yes julia is not a good choice but it's perfectly fine for soft realtime.
[1] https://github.com/JuliaLinearAlgebra/Octavian.jl/issues/24#...
There's various options depending on your problem. You could just hold firmly onto memory, preallocating all your buffers before the program starts and then don't allow them to be GC'd, you could also use stack allocated arrays like in StrideArrays.jl / StaticArrays.jl, or some combination of the above. You could also just manually use ccall to malloc and free memory as needed I guess.
There's a nice talk here [1] about using julia in soft-real-time for robotics.
I'd say the biggest impediment to real-time programming in julia currently is not the GC, but instead that the language has lots of optional optimizations that the compiler can choose to not perform if it thinks it'd be beneficial to do so, and those optimizations can change between minor versions. Hence, there's a lot of testing you'd need to do to make sure your code really is doing exactly what you want and you won't hit the GC or dynamic dispatch, and you'll need to redo that testing every time you update julia or any packages, which is obviously not ideal.
Essentially most GCs require you to pay for it even if you don't use it in critical sections. I'm not sure if Julia has the semantics for explicitly disabling GC during hot loops, but it is extremely difficult to do that without the dedicate hooks.
The issue is not performance - GCs are usually faster than not. It's determinism.
For that, I'd probably just do multiprocessing and have one realtime process that only operates on preallocated or stack allocated buffers with the GC disabled, and then another separate process that does your latency insensitive stuff.
I'm sure other languages are able to handle this sort of thing more elegantly in one process with tasks, but it can be made to work in julia, and there are lots of situations like robotics where you may want to leverage Julia's excellent ecosystem in a realtime system, so I just disagree with your earlier statement about julia being incompatible with the application domain, even if it's not a perfect language for it.
May be one day, assuming the rust mentality becomes more mainstream but computational science has always been its own thing that doesn't really follow cs trends (both good and bad in its own respects) and it certainly isn't moving in the direction of more intelligent memory management.
As an aside, I don't see why Rust couldn't seep into "anything seriously academic". C++ is widely used, and I've seen a gradual uptick in the usage of RAII. Rust seems like the next step, especially with the effort put into the ecosystem and learning materials. Sure, it might come slower than in domains where safe memory management is critical, but I don't think the needs of academics are fundamentally different here.
FTR, I think it's fair to question whether numerical computing should have an outsized influence on the direction of the language. I also think it's a pretty fair comparison to point out how standardized and consistent the Rust governance process is compared to Julia's (the Rust RFC system is an exemplar here). That doesn't mean there is a dearth of PL and systems knowledge in the Julia community though.
[1] https://github.com/AlgebraicJulia/Catlab.jl [2] https://www.youtube.com/watch?v=tRBl-6uEJJE [3] https://github.com/JuliaParallel/Dagger.jl/ [4] https://github.com/FluxML/IRTools.jl, https://github.com/FluxML/MacroTools.jl [5] https://github.com/JuliaCompilerPlugins [6] https://github.com/0x0f0f0f/Metatheory.jl [7] https://2020.splashcon.org/details/splash-2020-rebase/13/Non... [8] https://github.com/JuliaSymbolics/Symbolics.jl
How many people with significant prior language design experience are deeply involved in the language design process? what about database dev experience? OS dev? server dev?
> I also think it's a pretty fair comparison to point out how standardized and consistent the Rust governance process is compared to Julia's (the Rust RFC system is an exemplar here).
Yep. I want the language to be good, but it needs help from communities like the Rust language team.
I know their goal is to really be an all around good general language, so what you said makes sense for that goal regardless. I'm just curious if you have insight into gains that the numerics people like me could get that we might not know enough to know we're missing.
Many Base functions are currently untyped, which means it's hard to abstract out the key properties that enable reuse, compiler optimizations, and parallelism. See [0] for a language expert's take on numerical computing, which involves defining traits like associativity that enable automatic parallelism by default.
A stronger culture of functional programming could make your code faster and easier to understand. It's often (though not always) easier for tooling to optimize pure functions over immutable data structures (and arrays). Arrays are currently mutable, and there is too much emphasis on mutating functions.
The iteration protocol can be made more memory-efficient for large collections and simpler, following Rust's implementation [1].
Macros are useful for high-performance computing [4], and Julia's macros can be made more composable [2].
In general, systems languages don't make language decisions lightly. They have committees, discuss how other languages do things, make proposals. This allows more perspectives on each decision. That would be an improvement over the more ad-hoc style of Julia development, as long as Julia can avoid adding every possible feature, which is a risk of expanding the decision-making body [3].
[0] https://www.youtube.com/watch?v=EZD3Scuv02g
[1] https://mikeinnes.github.io/2020/06/04/iterate.html
[2] https://github.com/JuliaLang/julia/issues/37691
For example base functions being untyped is exactly how good Julia library code must look. Julia wants your code to be generic. It will get specialized (and compiled) when called with concrete types. The type system is not designed to encode invariants about the library, but to pass through information from the call site to the compiler when encountering inner functions. From a design perspective it is duck typed.
A fundamental difference to compiled languages is that the compiler doesn't need to reason about the types ahead of time because it only gets triggered when the function is called, at which time the concrete type information is available.
From my perspective this interaction of parametric type system, JAOT compilation, and multiple dispatch/ubiquitous generic code looks like they are interacting in a way that is genuinely new and exciting. At least I don't know any language that does something similar.
One way to think of it is that Julia only has template functions, with many the draw backs and problems that entails. C++ 20 introduced Concepts to improve this aspects of the language, and I really believe Julia is in need of something along those lines, but this is more for humans than for the compiler.
And obviously Julia is doing very well when it comes to enabling reuse and composability for example. I mean, I can throw a Neural Network into a Differential Equation, solve it using a state of the art solver, differentiate through the whole thing to do a gradient descent, and run all that on the GPU or CPU with the same code. So it's pretty absurd to claim that the duck typing in the base library is a problem for reuse or parallelism.
BTW we also prototyped our problem space with Fortran and Julia, and Julia actually ended up faster than the Fortran implementation for the same algorithms. So compiler optimization also is not constrained in this way.
Finally, Rust is a language built on the principle to not look at the cutting edge of PL research but instead to look at established things and implement them in a sound and relatively conservative way.
I use and like the language.
> For example base functions being untyped is exactly how good Julia library code must look. Julia wants your code to be generic.
I don't think there needs to be a conflict between being typed and being generic, especially when traits are available. There has been a lot of advancement in flexible type systems and ad-hoc polymorphism in the last 20 years. For example, there's a lot we can say about the type of `map` or `filter`, but none of that information is specified in Julia.
Map should have type
map(f : T -> U, arr : Iterable{T}) -> Iterable{U}
None of that can be expressed in the Julia type system as is. There is no AbstractCallable type that functions would be subtypes of, and there is no Iterable trait (AbstractArray comes closest but without multiple inheritance of AbstractTypes you can not rely on it).> A stronger culture of functional programming could make your code faster and easier to understand. It's often (though not always) easier for tooling to optimize pure functions over immutable data structures (and arrays). Arrays are currently mutable, and there is too much emphasis on mutating functions.
Yes and no. Immutable data structures shine in 2 scenarios:
1. Small types that can be represented with a couple of machine words. Julia supports this already through via stack allocated immutable struct types and packages like StaticArrays [1] 2. Persistent data structures as found in most FP languages.
Note how neither of these capture the large, (semi-)contiguous array types used for most numerical computing. These arrays are only "easier to optimize" if one has a Sufficiently Smart Compiler to work with. Here we don't even need to talk about Julia: the reason even Numba kernels in Python land are written in a mutating style is because such a compiler does not exist. You may be able to define something for a limited subset of programs like TensorFlow does, but the moment you step outside that small closed world you're back to needing mutation and loops to get a reasonable level of performance. What's more, the fancy ML graph compiler (as well as Numpy and vectorized R) is dispatching to C++/Fortran/CUDA routines that, not surprisingly, are also loop-heavy and mutating.
Should Julia do a better job of trying to optimize non-mutating array operations? Most definitely. Is this a hard problem that has consumed untold FAANG developer hours [2] and spawned an entire LLVM subproject [3] to address it? Also yes.
> The iteration protocol can be made more memory-efficient for large collections and simpler...
Yup, this has been a consistent bugbear of the core team as well. The JuliaFolds ecosystem [4] offers a compelling alternative with fusion, automatic parallelism, etc. in line with that blog post (which, I should note, is a much different beast from Rust's iterator interface/Rayon), but it doesn't seem like the API will be changing until a breaking language release is planned.
> In general, systems languages don't make language decisions lightly. They have committees, discuss how other languages do things, make proposals. This allows more perspectives on each decision. That would be an improvement over the more ad-hoc style of Julia development, as long as Julia can avoid adding every possible feature, which is a risk of expanding the decision-making body.
I'd argue this is a property of mature, widely used languages instead of systems languages. Python, Ruby, JS, PHP, C# and Java are all examples of "non-systems" languages that do everything you list, while Nim and Zig (note: both less well adopted) are examples of "systems" languages that don't have such a formalized governance model.
Julia (along with Elixir) are somewhere in between: All design talk and decision making is public and relatively centralized on GitHub issues. There is no fixed RFC template, but proposals go through a lot of scrutiny from both the core team and community, as well as at least one round of a formal triage (run by the core team, but open to all). Any changes are also tested for backwards compat via PkgEval, which works much like Crater in Rust. There was a brief effort to get more structured RFCs [5], but I think it failed because the community just isn't large enough yet. Note how all the languages with a process like this are a) large, and b) developed it organically as the userbase grew. In other words, you'll probably see something similar pop up when the time savings provided by a more structured/formal process outweighs the overhead of additional formalization.
[1] https://github.com/JuliaArrays/StaticArrays.jl [2] https://www.tensorflow.org/xla, https://tvm.apache.org, https://github.com/pytorch/glow, etc etc etc. [3] https://mlir.llvm.org/ [4] https://github.com/JuliaFolds [5] https://github.com/JuliaLang/Juleps
I think this is an idealized description. From the iteration blog post: "While one of many issues on multi-line comment syntax has 121 comments, the iteration overhaul proposal just says ‘we hashed it out at JuliaCon’ and there’s no mention of alternatives or tradeoffs considered."
> you'll probably see something similar pop up when the time savings provided by a more structured/formal process outweighs the overhead of additional formalization
Imho time savings aren't the main benefit of formal proposals. Rather, the decision outcomes are improved.
It's not uncommon for good Julia codes to smoke Fortran codebases, like the DiffEq verse or with staged programming, see Steven G. Johnson's keynote at JuliaCon 2019:
EDIT: Expounding on "why": On what aspects specifically does the Julia community need help from system programmers and programming language experts?
If there's one thing Rust is not, it's a good "glue" language. It seems to do best in large, densely coupled projects. This is where there's the least cost to defining all your own types, and also where a strong type system provides the most benefit. I think there will always be a place for more dynamic languages that tend to do better at interfaces.