How Julia ODE Solve Compile Time Was Reduced from 30 Seconds to 0.1
sciml.ai
sciml.ai
It'll be interesting to see, going forward, how much of these recent tooling improvements will trickle down to more casual devs, such that they too, can get significantly lower compile times.
I'm pretty optimistic on that front. It looks like much of the improvements came from writing more inferrible code, removing code that breaks reasonable invariants in Base, using SnoopPrecompile, and using a sysimage.
The first two things should be done anyway, and better tooling will help with that. Future native code caching will probably get results similar to current sysimages. That leaves SnoopPrecompile, which is a small maintenance burden, but is a pretty minor addition in most case.
One approach is to include a block of code which exercises the functionality of your package. This block can then, at the same time be used to: 1. Check your code is inferrible 2. Run static analysis to catch bugs 3. Use SnoopPrecompile
However, it can be incredibly frustrating sometimes. This article is a great example of the kind of wizardry you have to do to figure out why something isn't as fast as you thought it would be. There seem to be a lot of, "If you don't do it this way, it'll still work, but it will slow as hell," practices that are in the community's collective knowledge, but are big stumbling blocks in the beginning. (making sure you have concrete types in struct fields, the Val type and dispatch vs an if statement, how to write "type stable" code, etc.)
I had a lot of incorrect ideas and assumptions about the type system and the specialization mechanisms until I watched the Julia internals talk on YouTube from like 2014. This is probably due to it still being relatively young, and it has gotten better over time.
There are a lot of things to like about Julia that make this bearable though! It's really a great language to work in.
The one that gets me most often is the boxing of captured values. I come from Scheme/Racket and I'm just so used to exploiting closures. I also struggle with "over-typing" function declarations. I'm getting better about that though, now that I have a better handle on what exactly the dispatch mechanism is doing.
Much of this is more complex than this. Most open source software doesn’t ship with any assumptions about a particular BLAS/LAPACK implementation at all - and on HPC systems you are generally expected to choose one as appropriate and compile your code against it. It is generally only when you download a precompiled version that you’re given a particular implementation, but it doesn’t mean you can’t use another one if you compile from source as the BLAS and LAPACK libraries just present a standard API. Generally, for performance reasons, you want to compile specifically for your platform, because precompiled wheels from Conda, PyPI, etc. will leave performance on the table.
On forward thinking cluster teams these days, sysadmins use tools like Spack and Easybuild and to some degree software is made available to available to users either directly or by request, so it’s usual to log into a cluster and have multiple implementations available to choose from and compile your code against. More often than not, it’s still on you however to compile against what you need as dependencies. It’s a worthwhile exercise in HPC to try different ones and check the performance characteristics of your code on the particular machine with multiple implementations.
I took a break from Julia a year or two ago because of some of these issues, one of the big ones being I didn't want to write and maintain a set of non-allocating LAPACK wrappers for iterative solvers, but the memory churn was killing my performance. So, so glad FastLapackInterface and LinerarSolve are a thing now, and the MKL situation is much easier with this trampoline development, makes me want to start working on Julia solvers again.
It does feel difficult to write performant Julia if you don't put a lot of effort to stay "in the know" as a lot of this knowledge is very dispersed, but I guess it makes sense as the language is still changing quite rapidly.