Potential of the Julia programming language for high energy physics computing
arxiv.org
arxiv.org
A couple of years ago, when I was so faschinated by different programming languages, I stumpled upon Julia. I don't exactly remember how but I think I was scrolling through a list of programming languages that looked interesting so I could write "Hello, world" in it.
So I stumpled upon Julia, knowing nothing about it and downloaded it. At the time the language was still very young, but I remember thinking to myself that this language has a lot of potential, it promises something we actually need, like being dynamic yet high performance.
I've slowly been following Julias progress overtime and damn, it is growing rapidly and gained a lot of attention! I'm happy because it deserves it.
I've been working with Matlab again and I'm thinking about giving Julia another shot, this is also because Apache Octave was a disaster for me.
The end.
Back when I was on ATLAS we mostly did things in C++. If I never have to do another matrix multiplication in C++ again it will be too soon.
Fortran would be less painful than C++, for math-heavy code.
I did some Fortran in undergrad, and while it's pleasant to work with, if you have to do anything beyond numerical calculations you are basically screwed.
C++ has the features as a language to be almost Fortran level pain for math heavy. But the libraries tend to fall short. E.g. Eigen kind of does it, but it will just explode with a thousand line compile error when you least expect it. Also it's missing some quite elemental stuff like ND arrays. And getting code to work across even minor Eigen versions is a crapshoot.
Xtensor seems to have some potential, although it was too buggy to use back when I last tried it.
However, due to many fundamental design mistakes in C++ I'm not really expecting any code, including math heavy, to be very plesant or productive in it.
Aside from "why not Fortran" the "why not 2 languages" is because moving your implimentation to a language few of your users know creates a big barrier between users and developers of your code. In Python, ~90% of users don't even know the language that the packages they use are written in which makes it a lot harder for them to become contributors. Using a single language means that as users learn how to use libraries they are also learning how to contribute to them in the future.
Sadly Julia is pain too if you don't do REPL/notebook (which you shouldn't). Julia has the design to solve the two language problem but not the implementation. And will probably never have because Julia community refuses to see this as a problem.
High level naturally pushes you towards abstract types, GC, not caring about allocation, do what I mean not what I said.
Low level naturally pushes you towards concrete types, deterministic mallocs, do what I said not what I mean.
e.g. do I want integers like Python or integers like C? Yes.
An interesting thought experiment is to think if someone was to release F# today as a new language but calling it something else, like Flosure or something. I bet people would lose their minds over it.
Finance programmers would go nuts!
MS did include lambdas as a primitive but also decided to use python as new formula language.
Saddly i think even this boat has sailed since they added python to Excel now. A true miss opportunity.
While this is sentimental, as in it can't be exhaustively explained by logic and known facts, I still think it's somewhat rational. Fool me once etc.
The newer generations probably don't really grasp what MSFT pulled on OSS, developers and computing in general to make it the monopoly it is. It's possible that it has changed course, but I'll probably never trust it has.
I want integers like Python that are as fast as integers like C. I don't care how they are allocated or deallocated. Javascript comes very close even without having integers at all.
Probably nobody expects to squeeze every last cycle with a higher level language. I'd say getting consistently within magnitude of C performance can be said to have solved the two-language problem. GC pauses etc are perfectly acceptable for e.g. data analysis.
Jupyter got rid of their direct javascript interface for security reasons that I don't agree with. Does Pluto.jl have a model where trusting display code is as easy as trusting the code being run?
Although I wouldn't be surprised them being faster than ROOT with e.g. asm.js.
JS+ASM greatest advantage is cross-platform which is one of the least features requested in the field. i.e No one expect you to be able to run your analysis code on windows.
My point is that JS is really fast for some things, at least compared to what some people's impressions seem to be. Even though it's very optimization unfriendly in its semantics. And that this hints that the two language problem is indeed solvable for analysis at least.
"Laypeople" may also think that code is optimized to the last cycle in something like HEP simulations. It's made fast enough and the optimization is nowhere near the level of e.g. graphics heavy games.
Real-time usage like high frequency large data collection will probably never happen on the "single language". But I'd guess ROOT is not used at that level either? Also at least last time I checked, ROOT is moving to Python (probably not for the hottest loops of the simulation though).
(Off-topic: C++ interpretation like done in ROOT seems like a really bad idea.)
Sorry for assuming that. I really felt the pain of thinking of possibility of combining two things I hate so much together (JS+ROOT)
> "Laypeople" may also think that code is optimized to the last cycle in something like HEP simulations. It's made fast enough and the optimization is nowhere near the level of e.g. graphics heavy games.
I understand that in other areas there might be more sophisticated optimizations, but does not change things much inside HEP field community. And it is not optimized only for simulations but for other things too. It is not one problem optimization.
> Real-time usage like high frequency large data collection will probably never happen on the "single language". But I'd guess ROOT is not used at that level either? Also at least last time I checked, ROOT is moving to Python (probably not for the hottest loops of the simulation though).
I did not mean to indicate that ROOT is being used to handle the online processing (In HEP terms). It is usually handled via optimized C++ compiled code. My idea is that you will probably never use JS or any interpreted language (or anything other than C++ to be pessimistic) for that. ROOT at the end of the day is much closer to C++ than anything else. So learning curve wouldn't be that much if you come with some C++ knowledge initially.
> Also at least last time I checked, ROOT is moving to Python (probably not for the hottest loops of the simulation though).
I think you mean PyROOT [1]? This is the official python ROOT interface It provides a set of Python bindings to the ROOT C++ libraries, allowing Python scripts to interact directly with ROOT classes and methods as if they were native Python. But that does not represent and re-writing. It makes things easier for end users who are doing analysis though, while be efficient in terms of performance, especially for operations that are heavily optimized in ROOT.
There is also uproot [2] which is a purely Python-based reader and writer of ROOT files. It is not a part of the official ROOT project and does not depend on the ROOT libraries. Instead, uproot re-implements the I/O functionalities of ROOT in Python. However, it does not provide an interface to the full range of ROOT functionalities. It is particularly useful for integrating ROOT data into a Python-based data analysis pipeline, where libraries like NumPy, SciPy, Matplotlib, and Pandas ..etc are used.
> Off-topic: C++ interpretation like done in ROOT seems like a really bad idea.)
I will agree with you. But to be fair the purpose of ROOT is interactive data analysis but over the decades a lot of things gets added, and many experiments had their own soft forks and things started to get very messy quickly. So that there is no much inertia to fix problems and introduce improvements.
The alternative being OS X, or maybe Solaris, if one was in a building that still had a couple of pizza boxes that weren't yet on the throw away pile.
Julia has BigInt and Cint. Maybe there could be a more Pythonic implementation that scales between machine types and BigInt. That would not be hard to implement in Julia. I just have not found a good use case.
What I like about Julia is that I can do both the high level and low level in Julia. Abstract code becomes concrete due to late binding.
You can access Libc.malloc and Libc.free if needed. See GPUCompiler.jl
The effect system exists. If you can prove no side effects, eager finalization is also possible. That is the finalizer and deallocation will run deterministically. It's new though so effect analysis is a mostly manual affair at the moment.
Can you be more precise on this? In my case, I do a lot work around visual outputs (such as images). So notebook based development (such as Pluto.jl) feels perfectly fine.
Usage where you don't keep the Julia process running. E.g. "$ julia stuff.jl" form a shell.
Do standalone binaries here include shared libraries (with a C ABI)? That would be a dream.
Unfortunately, the core devs are not too chatty about standalone binaries, because of how Julia's internals are set there are going to be a lot of unforeseen challenges, so they are not trying to promise how things will be rather let's wait and see how things will turnout. Since packagecompiler.jl already has C ABI and one goal discussed about binaries being easily callable from other languages and vice versa, I would bet that it will have shared libraries.
The discussion seems to be too deep in Julia internals for me to follow. Is this about startup time or defining an entry point (or both?). I haven't had problems with Julia entrypoints (yet at least).
With the nightly `./julia-f7618602d4/bin/julia -e "using DynamicalSystems"` still takes over 5 seconds. Can I somehow define a main to make this faster or precompile more efficiently?
> Since packagecompiler.jl already has C ABI and one goal discussed about binaries being easily callable from other languages and vice versa, I would bet that it will have shared libraries.
Sounds promising. Shared libraries are not a musthave for me, but could allow Julia save us from C++ in more cases.
DynamicalSystems looks like a heavy project. I don't think you can do much more on your own. There have been recent features in 1.10 that lets you just use the portion you need (just a weak dependency), and there is precompiletools.jl but these are on your side.
You can also look into https://github.com/dmolina/DaemonMode.jl for running a Julia process in the background and do your stuff in the shell without startup time until the standalone binaries are there.
Compiler latency ("time to first plot") used to be miserable but after a few releases with incremental improvements it feels mostly solved to me.
Just now on Friday at JuliaCon Local in Eindhoven one of the keynotes was about similar ongoing work on stand-alone binaries (including shared libraries to call like C/Fortran.)
https://julialang.github.io/PackageCompiler.jl/stable/libs.h...
A Pluto.jl notebook is a human readable Julia source file. The Pluto.jl package is itself developed via Pluto.jl notebooks.
https://github.com/fonsp/Pluto.jl
Also, the VSCode Julia plugin tooling has really expanded in functionality and usability for me in the past year. The integrated debugging took some work to setup, but is fast enough to drop into a local frame.
https://code.visualstudio.com/docs/languages/julia
Julia is the first language I have achieved full life cycle integration between exploratory code to sharable package. It even runs quite well on my Android. 2023 is the first year I was able to solve a differential equation or render a 3D surface from a calculated mesh with the hardware in my pocket.
How is the module reloading story nowadays? Last time I gave up the state of the art was Revise.jl, which was a world of pain.
Yes it does, unless you do crazy stuff (which you probably should put in a Julia file and let it precompile). The order of cells and execution order does not matter (unless you do crazy stuff).
Yes, you should give it a go (again)!
Is that via webassembly or... ?
julia> versioninfo()
Julia Version 1.9.3
Commit bed2cd540a1 (2023-08-24 14:43 UTC)
Build Info:
Official https://julialang.org/ release
Platform Info:
OS: Linux (aarch64-linux-gnu)
CPU: 6 × Cortex-A55
WORD_SIZE: 64
LIBM: libopenlibm
LLVM: libLLVM-14.0.6 (ORCJIT, cortex-a55)
Threads: 1 on 8 virtual coresAdding to that modular mojo, and i think we might have a repeat of DART vs typescript...
The big selling point from the getgo was that Julia is general purpose. It's now over 10 years and it's still about as general purpose as MATLAB (although on a totally different level design-wise).
A big smell from the getgo was 1-based indexing. I appreciate that it's more familiar from some branches of math and physics (and MATLAB), and it's something I could live with, but more or less all general purpose languages index from zero. And there's a good reason for that. And that makes interoperability trickier.
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
I frankly would be concerned if the main contributors were not focused.
Julia is the first unique language I've been excited about in decades, as it can often outperform C, C++, and numpy in several use-cases. The only downside is until the 170MB+ lib is cached by the kernel, it gets panned by BS perf stats due to initial i/o constrained load times on some platforms.
Good luck, and have a great day =)
Numpy is OK, but verbose and a bit of a pain.
Given that ML is becoming so important and Julia has ergonomic handling of numerical code, why don't they explicitly support it? Is it inertia?
ML on CUDA is very much part of Julia already, and Meta seems into that lib as well. =)
Compared to the CUDA dumpster fire at a cat food factory, it is often trivial and nearly transparent syntax for the users familiar with ML.
Really depends on the use-case =)
The conventional options are fairly well documented:
https://fluxml.ai/Flux.jl/stable/gpu/
The big mistake was Microsoft not having .NET cross-platform from the beginning and not pitching F# as a Python alternative.
To play devil advocate, i think when F# was out (and still being developed), the Zeitgeist was really enamored with the idea of dynamic language, and the envisioned productivity gain from them. So F# was really too early to the dance.
Since then, things have change quite a bit with most major dynamic language adding a proto-type system.
Right now, i think we understand better what made language such as python and ruby so popular and are able to design better statically typed version with better tooling around them ( rust, zig, julia etc...)
I wonder if GPT-4 could be used to efficiently start porting more stuff over to Julia. I feel like Python is really a poor fit for deep learning. Julia is much more pleasant.
So yeah, I'd be surprised if FB ever moved away from Python.
There is Zygote.jl, which is used in Flux.jl. However, it's more in maintenance mode. At some point, Diffractor.jl was hyped but it didn't take off yet. And then there is Enzyme.jl which people hype now.
But for me as a user, it's not clear what I should really do to make my code well differentiable for those libraries.
If you stick with torch, jax or tensorflow, everything seems to work better regarding AD.
[0]: https://github.com/FluxML/Flux.jl [1]: https://github.com/FluxML/Zygote.jl [2]: https://github.com/JuliaDiff/Diffractor.jl [3]: https://github.com/EnzymeAD/Enzyme.jl
It seems to me that Julia is still moving too fast for its own good. At some point, something needs to stabilize for it to be worth adopting.
People often conflate popularity with good design. Have a great Monday =)
Python is the core issue rather than specifically Numpy, as it was bodged on (SWIG) to try to fix real use-cases... much like how 30 years of bodged on GPU mailbox structures made people biased to implement ridiculous solutions that necessitated ridiculous software paradigms.
I don't think any one person has the will and resources to resolve the core problem. However, Julia will likely also initially bind the same nonsense for awhile due to compatibility needs, but users tend to re-factor nonsense out of the ecosystem with better native solutions over time.
"All software is terrible, but some of it is useful" =)
It seems to me that, even given Python's constraints as a language, that a nicer wrapper could have been developed in Numpy that still called out to bindings to do the heavy work.
> functional compatibility layer for a fundamentally broken language paradigm
What did you mean by this? Is it mainly the fact that Python is a bit an undesigned language in general that calls into unmanaged languages like C/C++ and Fortran in its libraries?
Why would one bottleneck a design with polyglot stacks even before a single line of code was implemented? This sounds like naive nonsense.
"What did you mean by this? ... Python"
Python was never designed to handle threads or parallelism properly, and has performance issues Numpy tries to address though its wrapped C/C++ libraries.
Python became 30 years of spiral development, and implodes into a new implementation every so often. Depending on the use-case it may prove appropriate, but never optimal. =)
https://gist.github.com/AlexanderFabisch/6343090 (small feature comparison list, but ignore the incorrect opinions)
https://fluxml.ai/Flux.jl/stable/gpu/
https://alan-turing-institute.github.io/MLJ.jl/dev/about_mlj...
You don't seem to know much about what you're criticizing.
Just remember to put the value in quotes when initializing:
BigFloat("1.0") NOT BigFloat(1.0). Without quotes it's silently initialized as double precision (!)
julia> BigFloat(1.0) |> typeof
BigFloat julia> BigFloat(1.1)
1.100000000000000088817841970012523233890533447265625
I'd recommend using the big string macro though isntead of converting explictly from a string: julia> big"1.1"
1.100000000000000000000000000000000000000000000000000000000000000000000000000003(And, in Julia’s defense, this is all spelled out in the documentation, immediately accessible from the REPL. I just never read it carefully enough.)
My concern is most physicists won't read that either and may not even notice the loss of precision until it's too late (after a paper has been published).
With PackageCompiler.jl [6] you can even make AOT compiled standalone binaries, though these are rather large. They've shrunk a fair amount in recent releases, but they're still a lot of low hanging fruit to make the compiled binaries smaller, and some manual work you can do like removing LLVM and filtering stdlibs when they're not needed.
Work is also happening on a more stable / mature system that acts like StaticCompiler.jl [7] except provided by the base language and people who are more experienced in the compiler (i.e. not a janky prototype)
[1] https://docs.julialang.org/en/v1/manual/embedding/
[2] https://pypi.org/project/juliacall/
[3] https://www.rdocumentation.org/packages/JuliaCall/
[4] https://github.com/Clemapfel/jluna
[5] https://github.com/Taaitaaiger/jlrs
This seems false to me. StaticCompiler.jl [1] puts in their limitations that "GC-tracked allocations and global variables do not work with compile_executable or compile_shlib. This has some interesting consequences, including that all functions within the function you want to compile must either be inlined or return only native types (otherwise Julia would have to allocate a place to put the results, which will fail)." In practice, this means that you can therefore not use the base library in your own programs. PackageCompiler.jl [2] has the same limitations if I'm not mistaken. So then you have to fall back to distributing the Julia "binary" with a full Julia runtime, which is pretty heavy. There are some packages, such as PySR [3], which do this. It seems pretty usable for research I'd say, but difficult to put in production.
There is some word going around though that there is an even better static compiler in the making, but as long as that one is not publicly available I'd say that Julia cannot easily be called from other languages.
[1]: https://github.com/tshort/StaticCompiler.jl
Calling julia from another language does not in general require AOT compilation, though it does help to make things more portable and self contained. If one wants, they can literally just spawn a julia process, define a function on the fly and then hook into that process and call that function from another language.
> PackageCompiler.jl [2] has the same limitations if I'm not mistaken
You are mistaken. PackageCompiler works by basically bundling an entire julia system image with your program and the whole compiler, runtime, and stdlib compiled into one bundle together. If a piece of code works in the REPL it'll work from PackageCompiler
Okay but how do you distribute the program that depends on Julia to clients? You then also need to ensure that they have the right Julia available on their system. That's much heavier than dynamic binaries (which rely only on some shared libraries). My point was that it's not always "easy" to call Julia. That it's easy in HPC is a reasonable statement since you probably have full control over the system and it has loads of RAM available, but I wouldn't call it easy in general.
> You are mistaken.
I guess we're both right. I was talking about the "libraries" functionality [1].
[1]: https://julialang.github.io/PackageCompiler.jl/dev/libs.html
That's not the question that was asked though. The question the person asked as "Can you call Julia routines from other languages?" and the answer to that is "yes, you can easily do that".
__________________________________________________
> I guess we're both right. I was talking about the "libraries" functionality
I'm confused, what functionality exactly are you saying is missing from the libraries functionality? That very example you link to shows one using the julia runtime via printing.
And yes, dynamic dispatch and dynamic code generation do work from libraries created by PackageCompiler.jl as well...
That said, I think the bridges are actually pretty developed compared to what exists between many pairs of high-level languages. They are clearly working for some people since you will find Python-wrapper packages for Julia code. I'm not so sure it's so important to have a single `.so` package. There are solutions which allow Python to be-in-charge and bundle Julia to various extents. There are at least:
1. https://github.com/jlapeyre/julia_project
2. Usage of Conda and https://github.com/JuliaPy/pyjuliapkg
edit: just found https://docs.julialang.org/en/v1/manual/embedding/#Thread-sa...
Ok, so, it's tricky, but at least you can start an interpreter per thread...
It produces quite big binaries, but the more involved and complicated your application is, the less that matters.
The biggest selling point of Julia is that it's as easy to write as python if you don't care about performance, and if you do care about performance it's as fast as C++ if you can just follow some simple principles. This actually works in practice. For my day job I write the slow path exactly as I would python, and that's a really fast way to write code.
If you want a slow language that can call fast code, you already have python.
For me this is the biggest point for giving up on Julia. Julia has a lot going for it in the basic design and I was really excited about it in the beginning.
But even after all these years the environment is sadly almost unusable if you want to do something saner than REPL/notebook.
The TTFP is too damn high and I don't think it will ever be fixed.
I work in "one of those industries" where failure can easily cost millions in the space of minutes, and we use Julia almost exclusively. Unfortunately I can't give you details because of reasons, other than tell you you're very poorly informed.
> you're locked in to Julia
You're not "locked in" to Julia any more you're locked in to Python. Arguably less, as Julia can call C natively, Python can't.
My point isn't that you can't interop, and it seems you're intentionally (?) equivocating on this point. Surely you didn't misinterpret what I wrote as "you can't iterop"? My point is that you can more benefits if you don't. Like I said, if you want a slow language that can call fast routines, why not just use Python.
But I know from experience that this debate will not go anywhere. It's is the same old refusal to see the problem: "you're using it wrong".
Julia is dying and this is why it will die.
I vaguely recall a post from an ex Julia dev who had raised some issues regarding correctness of computations in Julia; I am not very familiar with the situation, have these concerns been addressed?
Edit: now that I recall better, I think the correctness concerns had to do with ambiguous composition of disparate components due to vague interfaces. I think this has to do with multiple dispatch in Julia, but I don't know details. I would love any info about this.
You'd have to be more specific. Currently ambiguous dispatch throws an exception. I don't think this is a recent change either.
You might have a function:
foo(x,y) = x + y
and multiple dispatch means you can call that with integers, or floats, or arrays, or GPU-resident arrays, or automatically-differentiating numbers, or symbolic algebra terms, or...
Just because you can, does that mean you should? It depends...
foo(x::Number) = x+1
And then someone else creates a new type thats a subtype of Number. And they run foo(x) and get errors or unexpected output.
Problem is the new type they created doesn't follow the assumptions that foo expects. Throw multiple dispatch into the mix and it gets even harder.
> Problem is the new type they created doesn't follow the assumptions that foo expects.
If you define that function and someone else (i.e. a user of your lib) defines a subtype of Number and calls your function on it, and it fails, then they haven't respected the interface to Number. There's nothing wrong with that function or type specialization in that case. In that case it's all about defining useful interfaces and respecting, as much of programming is.
E.g. if someone defines a subtype of Number because he has something that's kind of like a number in some respects but not others, maybe that shouldn't be an subtype of Number.
You also shouldn't define functions that have obscure conditions for being properly useable - at least if you expect your code to be useful more broadly. Nothing here is specific to Julia though.
If you have other specific examples we could discuss more.
using Unitful: m; foo(2.0m)
With the above definition, this will give DimensionError: 2.0 m and 1.0 are not dimensionally compatible.Probably this means you should define `foo(x::Number) = x + oneunit(x)` to respect the Number interface. But this interface isn't very strictly defined. I believe `Base.oneunit` was added after someone started writing the Unitful package -- building something useful in a legal grey area, and formalising later?
I don't know/use Unitful, but if that function didn't fail it's because the guys who wrote Unitful defined a promotion rule from Int to their Unit thingy. So.... that's how their type works. Don't use it. Or better, open an issue on github, they might have an explanation that's escaping us. I suspect it has to do with their mental model of what Unitful is supposed to achieve.
Let me tell you that my intuition agrees with you. As an ex-physicist, if I was designing a lib called Unitful I wouldn't let you sum sum 1 meter plus 1 unitless thing.
EDIT:
Actually, I just tried running your code and I do get DimensionError
foo(2.0m)
ERROR: DimensionError: 2.0 m and 1.0 are not dimensionally compatible.
I'm guessing you have some seriously outdated versions of something?EDIT2: Sorry I misread you. You do get an error. Ok I see your point. Maybe the Julia docs should more explicit about what the Number interface entails. Is `+ 1` allowed? You're assuming the answer is obvious, but it's not to me. In particular, that's not generic at all. You're probably right about `+ oneunit(T)`
Which country are you in?
In any case, if the complaint is startup time, I do agree that this has been a bit of a thorn in the side. It also interacts poorly with a (in my view) pretty immature ecosystem for running services (for instance, the grpc library doesn't seem to be widely used or heavily maintained) or interacting with it via IPC. So it's great if you can spin up a notebook or repl and leave it running all day doing your work, or if you have a long-running batch process, but I agree that I've found it a bit tricky to interact with it from the outside world in the "usual" ways.
Having said that, it does work - for instance, a system I work on spins up a long-lived julia thread from within a python service using the very nice shared library, and calls into julia functionality through that thread - I've just found it to be tricky; but IMO still very much worth it.
Nowadays also called TTFX where X can refer to variety of things.
In any case, yep, makes sense, thanks!
I guess for the typical scientific experimentation with plotting use case, why not just use the repl / notebook approach that (I think?) doesn't really suffer from this issue, as it's already up and running? My experience with startup time being problematic is when trying to "productionize" things that are not within a scientific / plotting loop. But my impression has been that it works well when doing things that are. Am I missing it?
My scientific/plotting loop runs from Bash (something not solved by Pluto.jl). This allows me to interoperate with a lot of software that's not in Python. I can version the code with e.g. git and the analysis pipeline tends to get refactored into libraries and CLI tools in the process. With Julia this is in practice impossible.
I do wonder if Pluto.jl might solve a lot of these issues for you though. I haven't really used it much, but it seems like it works with just plain .jl files, so you can non-jankily version them (unlike .ipynb), and extract libraries and refactor in all the normal ways.
My frustrations trying to do normal-software-development-principles with julia are actually more about the convention the community has around structuring files and modules, which just hasn't clicked for me at all and is a frequent source of confusion. (The testing framework / layout / something about it also hasn't clicked for me yet either, and I wonder if it just isn't great and that's why people don't seem to write many tests...)
But there's also just so much great stuff about the language that I'm not even really mad about any of this :)
That is more accurate.
> I do wonder if Pluto.jl might solve a lot of these issues for you though. I haven't really used it much, but it seems like it works with just plain .jl files, so you can non-jankily version them (unlike .ipynb), and extract libraries and refactor in all the normal ways.
It may solve enough issues at least for some cases. I'll check it out.
> My frustrations trying to do normal-software-development-principles with julia are actually more about the convention the community has around structuring files and modules, which just hasn't clicked for me at all and is a frequent source of confusion. (The testing framework / layout / something about it also hasn't clicked for me yet either, and I wonder if it just isn't great and that's why people don't seem to write many tests...)
From what I remember from the last time the module system was indeed quite a pain. The using/include/use seems like a total mess, and that IIRC Stefan Karpinski dismissed files-as-modules solely on the reason that Python does it that way did not really bring my hopes up.
I appreciate that the multiple dispatch design (which I find very much a pro for Julia) does bring about some additional challenges for a module system. But my hunch is that majority is because Julia was so strongly made by MATLABists to MATLABists (can replace MATLAB with e.g. R or Mathematica) that basic features required for a general purpose language were not much taken into account.
> But there's also just so much great stuff about the language that I'm not even really mad about any of this :)
This is exactly why I'm so mad about this! I would love to use all that great stuff AND follow good programming principles, but now I can't.
Time To First Plot. It was a meme back then that you needed to wait several minutes to draw just a simple line plot after starting up a Julia REPL/notebook because of compilation times (whereas in Python you're done in seconds with matplotlib.) Seems this has been reduced in Julia 1.9 with automatic pre-compilation of packages, but plotting still isn't "instant" as you would expect in scripting languages like Python.
Everybody doesn't use the REPL. In my, I think quite well founded, opinion nobody should use the REPL for more than one or two lines.
I don't find my plt.plot(x,y); plt.show() in a script ideal either. Something like Pluto.jl (given it's easy to modularize and preferably can be run without a notebook environment) is most likely better for a typical data analysis workflow (and will likely lead to more reproducible results and better tracking of result provenance).
For different cases, e.g. application development, the latency is not that critical. Although I find it quite insane that we accept so slow code iterations this day and age.
If the problem was just running a daemon, it would be an inconvenience at worst. The problem is that changes to the code tend to trigger a recompile, can't be made (e.g. structs can't be redefined if I recall correctly) and most crucially the session can get "silently" into some unexpected confused state.
If you don't need to change the code (and it's structured well enough not to hold a state between calls), then running in a process is no problem. But for that you can do just plain old AOT.
Aside from my personal gripes, which I have probably aired here more than enough for everybody, I think the design that more or less forces REPL/notebook style is detrimental to quality of scientific code. Scientists don't learn to modularize their code and don't learn to understand the basic control flow. I see this all the time with my colleagues and the "pedagogical problem" of notebooks/REPL is very common when I teach programming. E.g. a common problem is that students don't understand variable assignment because with REPL/notebook it doesn't necessarily follow the normal program flow.
I'm kind of amazed how little the scientific programming scene is concerned about this. From software development perspective (I'm an ex software developer turned scientist) having clearly defined state transitions, independent modules of code and the code doing what it says are really really fundamental things, without which you have no chance in hell in making anything more than a few lines to not turn into an unmanageable and buggy mess.
Maybe there's a (mis)conception that scientific code is somehow fundamental different beast than "general programming". But it's not. It's just a relic of having bad systems languages like C and C++ and bad special purpose languages like MATLAB and R. That Python can be and is used for both is an existence proof that there is no such divide. But we need a better Python for both general and scientific programming.
So, two things here: First of all, I have greatly appreciated the discussion, so don't feel like you've over-aired your thoughts!
But also, here's this, in my view, hyperbole again, I just don't see how "more or less forces" is the right way to describe the state of choosing whether or not to use a notebook/REPL style for julia. It works perfectly fine to develop with a different style. There's a little bit of startup time, but it's just not nearly a big enough deal to say it "forces" a different development process.
I think your criticisms of notebook driven programming make sense, but are also slightly overblown; in my experience, scientists are perfectly capable of writing good modular code, despite preferring to work in notebooks most of the time. It's not not a problem, I just don't think it's been a very big problem, in my experience.
But I very much agree with your last two paragraphs. I just don't think notebooks / REPLs are a big part of the reason things have ended up this way. I think it's more of a cultural thing that gets passed down from generation to generation, and it has to do with respect; I think scientists have never taught one another that code deserves respect and professionalism. I think they have always thought of it as "some throwaway stuff that generates the data for the paper"; the paper is the thing deserving of respect, not the tools used to create it.
But I think this has been slowly changing as more and more scientists are digitally native and learn increasingly more software skills as a first class concern.
More accurate is that the REPL/notebook style is currently so much more convenient to get started and to "explore" with that people will use that and it's what all Julia examples etc teach. Pluto.jl may be a good "transition" from this though, although I'm afraid modularization is still too inconvenient with it.
The script-workflow latency is also a reason why "scripters" don't switch to Julia even though they otherwise would.
So while the it's not practically impossible as I hyperboled, it is IMHO a blocker for wider Julia adoption and a transition to better programming practices.
In my experience in both teaching and helping working scientists with their code problems I think the current situation that REPL/notebook is so much easier to get started with is at least a major reason why better practices don't get adopted.
Scientists themselves doing any programming is quite a recent thing. Usually it was "lab engineers" doing "the coding" and data mangling and the scientist doing "the analysis" in something like SPSS or Excel or copypasted R oneliners.
It's also really understandable why scientists don't make much effort to learn and enforce good programming. It's a relatively minor part of their job, they manage to scrape much of the analysis together with poor practices, and when it almost inevitably explodes as the analysis gets more complicated, someone "more technical" gets called to fix the mess (I do this all the time).
In most languages modularizing code is made needlessly hard, and I'm smelling something like this in Julia module system (and in e.g. Python packaging). I sometimes wonder if there's some (unconscious) gatekeeping to keep "the using" and "the programming" (and "the core development") apart artificially. You actually sometimes encounter explicit statements like this (Linus Torvalds famously about C++ but in this post's discussion something like that not knowing R means that you shouldn't use ML (which is extra bizarre because very very few actually know R)).
The benefits of modularization are also not easy to foresee until you learn from many hard lessons of ending up with unmanageable sphagetti or analysis that is wrong but it may be even impossible to pin out why due to lack of reproducible code. And if you don't know about the ways of avoiding/mitigating this, you're not really equipped with learning the lesson.
When this compounds with significant hurdles to do "the right thing" I find it almost inevitable that "the right thing" will not get adopted.
As a sidenote I find it also quite odd why so little effort is done to "show the work" with data analysis when it's literally required when you do it in math class. I think they are almost the same thing. Maybe it stems from the computer seen as a calculator and "the work" is done on paper, but this is mostly due cultural lag from the era where computers were in practice calculators.
TTFP is (close to) solved. 1.9.0 TTFP is under 3 seconds for me, which I find about acceptable for script-type workflow. However, time-to-first-DynamicalSystems is still around 10 seconds, which is not.
I just installed and ran on my macbook - took me 7 minutes to get it installed and run through to a first plot, most of that was installing and compiling the Plots library.
Once it's installed and available I can run it in about 1 second and then execute
using Plots
function test_plot()
x = range(0, 10, length=100)
y = sin.(x)
plot(x, y)
end
@time test_plot()
which gives me a result of 0.048700s and a nice plot.
I don't think this is the defining issue with Julia now - it definitely used to be 10 years ago - but now it's more than competitive with Python in this regard.
The issue is really support and buy in for all the new cool libraries that folks make available to use.
$ time ./julia-1.9.4/bin/julia ttfp.jl
0.325806 seconds (202.80 k allocations: 13.208 MiB, 84.23% compilation time)
real 0m2.401s user 0m2.357s sys 0m0.553s
I.e. the TTFP is 2.4 seconds, not 0.33 seconds.
That said, 2.4s is huge progress (e.g. with Julia 1.6.0 it's 10 seconds for me) and acceptable for most simple plotting.
But just adding e.g. "using DynamicalSystems" brings this to 14 seconds, which is not acceptable for iterative editing. And it makes any Julia program using "DynamicalSystems" wildly impractical to use in a pipeline with other programs.
$ time ./julia-1.10.0-rc2/bin/julia -e "using Plots; using DynamicalSystems;"
0m5.433s
This is a different conversation from
julia> @time plot(x, y)
0.000465 seconds (484 allocations: 45.992 KiB)
But I agree that the conversation is pointless. The TTFX will never be fixed for Julia to be used with sane workflows because the few still using Julia don't want to accept the problem. But the flamewars are fun anyhow and maybe they prevent the next Julia from making the same mistakes.Wouldn't something like slow interpreter option for "glue code" be feasible? Most code doesn't need to be fast, and waiting for the "hot functions" to compile on change wouldn't be too bad. And you could still AOT for maximum speed when needed. This is essentially how Javascript JITs manage to start up so fast.
Another fine "hack" for many cases would be to keep the Julia process running (like a daemon) but making sure the state is clean when a script is run on it.
Something like PackageCompiler is workable, at least for the short term. Long compiles are OK if they don't have to be done all the time. I indeed manage to get < 1s TTFP with DynamicalSystems with it (last time I tried a while ago I didn't manage). The UX for creating sysimages could be nicer, but doesn't seem anything that a quick script can't fix. Maybe I'll give Julia another go in my next analysis.
From what I gather there's now a new approach for caching the compilation results. I really hope it will succeed, but I'm a bit jaded from the many "Julia 1.x fixes the TTFX" news.
The script typically loads the data files and runs simulations (a lot of times when their parameters are optimized). This means the models have to be fast, but they are also recurrent in nature so they can't be vectorized to numpy. So depending on the case it's usually numba, cython or C++, all of which can be quite painful.
Can't you compile the pipeline into a single script and then run it in a single instance of whatever?
Or at least semantics of that. The steps should be independently callable and each call should result in the same output with same parameters, i.e. the state would reset for each run. I don't care so much how this is exactly implemented as long as it fullfills these.
The data doesn't need to stay in the process memory. It's can be e.g. mmapped and is cached by the OS anyway. And deserializing tens or even hundreds of megabytes takes usually less time than Julia takes to import Plots.
This is the thing that isn't feasible with Julia due to the TTFP (each step incurs the startup latency).
> Wouldn't something like slow interpreter option for "glue code" be feasible? Most code doesn't need to be fast, and waiting for the "hot functions" to compile on change wouldn't be too bad. And you could still AOT for maximum speed when needed. This is essentially how Javascript JITs manage to start up so fast.
That is something I've been meaning to investigate more. A lot of the Plots.jl code is internally dynamic because of its design to hold everything in a Dict{Any} and recurse the plot recipes. In other words, it doesn't even benefit from having the JIT at all because internally everything is uninferred boxed variables. There's probably a good way to just do "please interpret the calls in this part of the pipeline" and not see any runtime difference but chop out the core of the compilation. It's relatively straightforward to do, it's a one liner using JuliaInterpreter.jl @interpret, and I think one dig that results in a win would be a nice example of what should be a pattern we do more often in the near future.
For now the interpreter seems to focus on debugging, and if I understand correctly, it can't be really used to e.g. speed up imports or to specify hot code to be AOT'd/JITed?
Do you have anything that you've shared publicly?
Codewise it's not pretty, especially the data mangling, but other people have managed to use it for their own purposes.
The model had to be implemented in C++ and wrapped for python bindings. It's a pain but sadly less painful than doing it with Julia's startup time.
Ok thanks for sharing.
(I still disagree about everything you said about Julia though. I bet you I could re-implement this faster in Julia thank your C++, and in a type generic way too which would make it trivial for users to plug in their types, and it would be trivial to package - I'm assuming you didn't package that repo, you just expect your users to be shuffling files around. You're not contributing to improving the reproducibility situation in the academia, let me tell you xD)
I didn't package it and the data mangling part has very little use outside this specific experimental design. Also Python's packaging is quite shitty.
The C++ code is totally standalone and can be git-pulled and compiled how you like. C++ packaging is totally abysmal.
As I said in the comment and the README the code is a mess (although a piece of art compared to the notebook shit out there). It's purpose is that the results can be reproduced from the raw data.
I would love to make my analysis code cleaner. But there's zero incentive for that and very little time. Nobody sadly cares about the code, and most scientists can't code for shit. This is probably why Julia was designed so that it forces your code to be shit.
Sorry but the tone, but you kinda set it.
JS does this by multi-phase JIT. Python by being dog slow in the execution and AOT by being dog slow in the compilation.
I think Julia's "JIT-AOT" may be a good approach (and perhaps almost necessary for Julia's dispatch) too but it seems to be very hard to make start up fast.
edit: oops, maybe PythonPlot is what i should use, not PyPlot. But when I try that I see: ┌ Warning: `PythonPlot` 1.0.3 is not compatible with this version of `Plots`. The declared compatibility is 1 - 1.0.2.[ and doesn't seem to be any faster
Other people may have different needs, different experiences, and different organizational or legacy constraints that makes it so that switching from python to julia is just not on the table any time soon. That's fine.
Julia and Python continue to both grow in the scientific computing space, and I think Julia is becoming more and more of a known name, and an accepted tool in the scientist's toolkit. It may not be the tool the majority of scientists use any time soon, but that's not really a problem.
I think focusing on providing significant value to the dev community in a way that's not disruptive to their current workflow is a better metric of success.
My 0.02 is that Julia is awesome... if I was starting a new numerical project from scratch, I'd pick Julia 100% unless there was a very compelling reason to use another language (e.g. to interact with a legacy system). There are some Julia libraries that are so insanely good, like SciML/DiffEq, that I use Julia just to use those libraries (because I'm not smart enough to re-implement that stuff in some other language).
That being said, Python has a sufficiently good ecosystem that it's a sane choice for scientific programming where all out speed (or multi-node parallel computing) is not the top priority. Python with Numpy/Numba, RAPIDS and/or JAX can be very fast and it is very easy to write concise and fast programs this way. The code quality of RAPIDS and JAX are very high too and documentation is solid.
I'd also offer that Chapel and Rust belong in the "HPC/super-computing capable" pantheon of programming languages.
But I can’t for the life of me understand how any language designer signed off on one based indexing.
It's been a long time i have used any language with one based indexing, but i don't remember it being particularly important in my day to day.
You have to keep track of which language you're using anyway to produce correct syntax...
When building data-structures you're often using small unsigned integers that use the complete value-space, (e.g. u8s and their bitsets in tries, u16s and blocks with that size in sparse vectors).
With 1 based indexing, you can't do that anymore without first having to cast the index into the next larger integer, adding 1, and then doing the access, which is a PITA, adds instructions, and messes with prefetching, while you're already working with difficult, hard to read, and heavily optimised stuff.