Julia and Mojo Mandelbrot Benchmark
discourse.julialang.org
discourse.julialang.org
A more interesting question is which version is more elegant, ‘obvious’ and maintainable. (Deeply familiar with both, but money is on Julia).
Yes, more than raw speed, what impresses me is that the version of code in [1] is already a few times faster than the Mojo code - because that's pretty basic Julia code that anyone with a little Julia experience could write, and maintain easily.
The later versions with LoopVectorization require more specialized knowledge , and get into the "how can we tune this particular benchmark" territory for me (I don't know how to evaluate the Mojo code in this regard as yet, how 'obvious' it would be to an everyday Mojo developer). So [1] is a more impressive demonstration of how an average developer can write very performant code in Julia.
[1] https://discourse.julialang.org/t/julia-mojo-mandelbrot-benc...
> what I find more interesting here is that the reasoning and code type applies to any composite type of a similar structure, thus we can use that to optimize other code completely unrelated to calculations with complex numbers.
It matters less whether super-optimized code can be written in Julia for this particular case (though there's some value in that too), the more important and telling part to me is that the language has features and tools that can easily be adopted to a general class of problems like this.
My humble two pennies is that Julia is missing the influencer factor: being endorsed by widely known entities that will attract the attention of both corporate eyes and the hordes of developers constantly looking for the next big thing.
Your money might be on Julia but $100mln was just placed on the Mojo/Modular bet...
- Slow startup times (e.g., time-to-first-plot) kill it's a appeal for scripting. For a long time, one got told that the "correct" way to use julia was in a notebook. Outside of that, nobody wanted to hear your complaints.
- Garbage collection kills it's appeal for realtime applications.
- The potential for new code paths to trigger JIT compilation presents similar issues for domains that care about latency. Yes, I know there is supposedly static compilation for julia, but as you can read in other comments here, that's still a half baked, brittle feature.
The second two points mean I still have the same two language problem I had with c++ and python. I'm still going to write my robotics algorithms in c++, so julia just becomes a glue language; but there's nothing that makes it more compelling that python for that use. This is especially true when you consider the sub-par tooling. For example, the lsp is written julia itself, so it suffers the same usability problems as TTFP : you won't start getting autocompletions for several minutes after opening a file. It is also insanely memory hungry to the extent that it's basically unusable on a laptop with 8gb of ram (on the other hand, I have no problem with clangd). Similarly, auto-formatting a 40 line file takes 5 seconds. The debugging and stacktrace story is similarly frustrating.
When you take all of this together, julia just doesn't seem worth it outside of very specific uses, e.g., long running large scale simulations where startup time is amortized away and aggregate throughput is more important than P99 latency.
You can do real-time applications just fine in Julia, just preallocate anything you need and avoid allocations in the hot loop, I am doing real-time stuff in Julia. There are some annoyances with the GC but nothing to stop you from doing real-time. There are robotics packages in Julia and they are old, there is a talk about it and compares it with c++(spoiler, developing in julia was both faster and easier and the results were faster).
I have been using two Julia sessions on an 8gb laptop constantly while developing, no problem. LSP loads fine and fast in vscode no problem there either.
The debugger in vscode is slow and most don't use it. There is a package for that. The big binaries are a problem and the focus is shifting there to solve that. Stacktrace will become much better in 1.10 but still needs better hints(there are plans for 1.11). In general, we need better onboarding documentation for newcomers to make their experience as smooth as possible.
I just timed vscode with the lsp. From the point I open a 40 line file of the lorenz attractor example, it takes 45 seconds until navigation within that same file works, and the lsp hogs 1 GB of memory. That's 5x the memory of clangd and 20x worse performance; hardly what I would consider a snappy experience.
I have no doubt that julia can be shoe-horned into realtime applications. But when I read threads like this [1], it's pretty clear that doing so amounts to a hack (e.g., people recommending that you somehow call all your functions to get them jited before the main loop actually starts). Even the mitigations you propose, i.e., pre-allocating everything, don't exploit any guarantees made by the language, so you're basically in cross-your-fingers and pray territory. I would never feel comfortable advocating for this in a commercial setting.
[1] https://discourse.julialang.org/t/julia-for-real-time-worrie...
Having a REPL open is not the same thing as a notebook, if you feel like that, cool I guess.
That thread is old and Julia can cache compiled code now from 1.9 and onward. However, it can not distribute the cached code(yet).
Writing the fastest possible real-time application in c/c++ has the same principles as in Julia. It's not as shoe-horned as you might believe.
When developing Julia, the developers chose some design decisions that affected the workflow of using the language. If it doesn't fit your needs that's cool, don't use it. If you are frustrated and like to try the language come to discourse, people are friendly.
I know, I'm always "holding it wrong". And that's the problem with julia.
> Having a REPL open is not the same thing as a notebook, if you feel like that, cool I guess.
Both workflows amortize the JIT times away by keeping an in-memory cache compiled code. This makes a lot of smaller scripting tasks untenable in julia. So people chose python instead. That means julia needs a massive advantage elsewhere if they are going to incorporate both languages into their project.
> When developing Julia, the developers chose some design decisions that affected the workflow of using the language. If it doesn't fit your needs that's cool, don't use it. If you are frustrated and like to try the language come to discourse, people are friendly.
This thread was about why julia hasn't seen wider adoption. It's my contention that the original design decisions are a one of the root causes of that.
Autocompletion in Julia is also just terrible and the tooling really is lacking compared to better funded languages. No harm in admitting that (when Julia had no working debugger some people were seriously arguing that you don't need one: Just think harder about your code! Let's please bury that attitude...)
It was very strange, but I don't think it says anything about the community, except that people have different opinions and preferences, like in any community.
> I have never seen anybody in the community say the correct way to use Julia is in a notebook.
patrick's comment is fully in the past tense for this part, and that was indeed a pretty common thing for a long while in the past. Especially pre-1.0, before Revise became mature and popular, an oft-recommended workflow was to use a notebook instead of an editor. Or it would come in the form of a reply to criticism about the slow startup or compilation latency - "the scientists who this language is made for use notebooks anyway, so it doesn't matter" was a common response, indirectly implying that that was the intended way to use the language.
On garbage collection and real-time applications, there is this [2] talk where ASML (the manufacturer of photolithography machines for TSMC) uses Julia for it. Basically it preallocates all memory needed before hand and turns off the garbage collector.
On the same more about real-time, if your call stack is all type stable [3], the you can be sure that after the first call (and subsequent compilation), the JAOT compiler won't be triggered.
About static compilation, there are two different approaches * PackageCompiler.jl [4], which is rally stable and used in production today. It has he downside of generating huge executables, but you can do work to trim them. There is still work to do on the size of them. * StaticCompiler.jl [5], which is still in the experimental phase. But it is far from being completely brittle. It does puts several restrictions on the chide you can write and compile with it, basically turning Julia in a static type language. But it had been successfully used to compile linkable libraries and executables.
Some of the concerns you have about usability in your third paragraph have been worked on with the 1.9 and 1.10 (coming) releases. The LSP usage is better thanks to native code caching, maybe you can try it again (of you have time). The debugging experience I honestly think is top notch if you're using Debugger.jl+Revise.jl [6] [7], still I know there are some caveats in it. About stack traces, there is also a more of work done to make them better and more readable, you can read the with done in these PR's [8] [9] [10] [11, for state of Julia talk].
Still, I can understand that Julia might not be able (yet) to cover all the usecases or workflows of different people.
[1]: https://julialang.org/blog/2023/04/julia-1.9-highlights/#cac... [2]: https://www.youtube.com/watch?v=EafTuyy7apY [3]: https://docs.julialang.org/en/v1/manual/performance-tips/#Wr... [4]: https://julialang.github.io/PackageCompiler.jl [5]: https://github.com/tshort/StaticCompiler.jl [6]: https://github.com/JuliaDebug/Debugger.jl [7]: https://timholy.github.io/Revise.jl/stable/ [8]: https://github.com/JuliaLang/julia/pull/49117 [9]: https://github.com/JuliaLang/julia/pull/49959 [10]: https://github.com/JuliaLang/julia/pull/45069 [11]: https://www.youtube.com/watch?v=2D8oRtDJEeg&t=2487
LLVM helps with its own optimizations, but more and more of those optimizations are being moved to the Julia side nowadays (since the language has more semantic understanding of the code and can do better optimizations). I believe the main thing LLVM helps with is portability across many platforms, without having to write individual backends for each one.
The gap will definitely narrow.
It's always fun though when one language does better naïvely in a benchmark to delve in and see how to match or surpass them, and see if it was worth the trouble.
Microbenchmark performance for languages in this class definitely shouldn't be seen as a strongly deciding factor though.
It's not only the actual benchmark numbers though. Its understanding the code that reaches those numbers: Can Julia do explicit SIMD? How awkward is that in one language or the other? Are there idiosyncratic bottlenecks? Bad design decisions that needs to be worked around in one language but not the other? And so on.
How the language gets to that number is of course very enlightening
Julia's poor AoT support (with small binaries) is a major Achilles heel. I really wish that the Julia developers had taken that more seriously earlier on.
Wasm fluid simulation in Julia: https://alexander-barth.github.io/FluidSimDemo-WebAssembly/
Differential equations demo in the browser: https://tshort.github.io/Lorenz-WebAssembly-Model.jl/
Someone even setup Julia to run on AVR mcus for Arduino (Directly using gpucompiler which statictools uses to compile to binaries):
https://seelengrab.github.io/articles/Running%20Julia%20bare...
Static compilation is indeed possible with Julia. But it's very limited in its capabilities and certainly not as effortless as a simple `mojo build myfile.mojo`.
From my personal experience. I've done graphical apps in GTK3 in Julia with PkgC.jl cross-compiling from Linux to Windows. And they worked. :)
* Massive executables (which you mentioned). This makes it very difficult to use with embedded systems.
* Functions are not precompiled by default. You need to write a precompile script [1], which leads to a "two script problem": one script to do what you actually want, and another script that (hopefully) hits all the types you'll possibly need at runtime. And yes you can use `--trace-compile=file.jl` or SnoopCompile.jl instead, but this is still another step I need to worry about when compiling something.
[1]: https://julialang.github.io/PackageCompiler.jl/dev/sysimages...
The problem is funding. There are 0 full-time employees working on this issue because JuliaComputing has gotten about 10% of the funding Mojo has.
You can imagine what a company like Boeing might be interested in when it comes to a programming language.
But we can only guess from the outside, and it's ultimately upto JuliaHub to decide how to spend the money, so I'll cross my fingers and hope that this gets us AoT static compilation sooner!
[1] https://jump.dev/ [2] https://julianlsolvers.github.io/Optim.jl/stable/
Connection to the Python ecosystem. Python remains the number 1 teaching language by a large margin.
AI funding. If they can get the buy in from the AI community that Julia never got, they have the resources to engineer around any challenges faced.
Solid foundation in modern language design, and with that, a focus on correct code produced by larger teams.
Yes, the need to run CPython interpreter is what makes Mojo slow (and it will remain that way, unless they abandon their "superset of Python" promise).
It's growing, but certainly not growing exponentially or anything like that. Here's some statistics from January this year: https://info.juliahub.com/julia-annual-growth-statistics-jan...
Regarding your negative experiences, the bad news is that we haven't solved those issues, but the good news is that we're making real progress on them. Version 1.9 released in may of this year and is the first version to cache native code from packages, which makes loading of julia code MUCH faster through more AOT compilation, and there are even more improvements coming in v1.10 later this year. https://julialang.org/blog/2023/04/julia-1.9-highlights/#cac...
Error messages are also receiving a fair amount of attention, but it's a hard problem and there's less agreement on what the best way forward is. However, there's been some good work going into improving the readability and clarity of error messages that I think will help alleviate these struggles.
I've used a bit for various pet projects (mainly some graph search and JuMP stuff) and it was convincing. But now I can see Fortran perform in real production code (where it shines, at the cost of being so antiquated that it's not funny anymore) and my expectations for Julia are now higher.
I'll give it another try 'cos you spend some time answering my question :-) (and because the charts in the 2nd provided link are just really convincing)
Let us know if you're experiencing any new painpoints too, or if things like code loading aren't as fast as you had hoped, there might be things we can do to help.
is it in the us, europe, asia?
I know cba and anz in australia supports julia internally for monte-carlo.
To a first approximation, the only people that love lisps are people with a solid computer science background, and most people working with AI and data processing day to day do not have a computer science background. They're scientists, engineers and mathematicians who see programming and programming languages as a tool needed to do their 'real' job and not as an end in itself. Python is the perfect language for people who want to learn as little programming as possible so that they can get on with what actually interests them.
If you mean “primarily uses S-expressions” then I guess I dont really see why that’s so important to you.
If you mean a language that is semantically similar to lisps and learned a lot of the important lessons that Lisp taught the programming world, I think Julia is one of the Lispiest languages in this space right now.
The syntax may not be S-expression based on the surface, but our Exprs are actually essentially just S-espressions so writing syntactic macros is very easy. The language is about as dynamic as is possible without major performance concessions, and is very heavily influenced by a lot of design ideas from the CLOS, with some features missing but also some cool features CLOS doesnt have.
I really hope I am wrong. I love Julia and would like to see it succeed everywhere, but it doesn’t seem to be happening.
Also, there’s nothing wrong with niches. Julia is undoubtedly less of a general-purpose language like Python, but it very much shines in its domain.
Can you elaborate some more on this? My worldview assumed that a lisp used a list as a primary code/data structure and Julia doesn't seem to be doing that... Of course it does provide a way to manipulate code and data because of its macros. But what makes a lisp a lisp?
"Code as data" is not quite the same thing as homoiconicity, which I feel is Julia's missing piece.
Julia is growing and evolving and finding new users and niches. It certainly hasn't 'won', but it's a bit early to call it a loss.
Julia is basically a Lisp under the hood. From playing around with it, it seems like the REPL experience is up there too.
The runtime/REPL is pretty good too, and can be quite dynamic with Revise.jl, but doesn't have the Real REPL-driven Programming Experience™ as defined here: https://mikelevins.github.io/posts/2020-12-18-repl-driven/
I'm not sure how much of the "breakloop" functionality Infiltrate.jl provides, but at least the runtime re-definition of types isn't supported in Julia, and is one of the shortcomings of the Revise.jl based workflow.
All this is not to take away from the original point, Julia does get you a big chunk of the way to being a Lisp and gives you a lot of expressive power. It's just to say that Julia is not just a reskinning of a Lisp with familiar syntax, it has some important design and implementation differences.
> I know I shouldn’t say so but I can’t help...
Remarking that the comment should not be taken too seriously, as it might be inappropriate.
Finally, saying the whole community is condescending given 1 in 32 comments is... a little rounding up from the statistics there.
I would say 1 in 32 is also about the experience. I stopped visiting Discourse, chatting in Slack because I found it exhausting that every time Python is mentioned someone came up with a different way of saying how much they hate Python. I know I wasn’t the only one disturbed by this but in the end communities make their own choices.
Sorry about missing the mojo specific [].
Some people feel that they should get to act like a prick, while everyone else should be humble and courteous.
If you act like a decent person person, there's no problem pointing out weak points in Julia and request help to work around it. If your only input is "Julia kinda sucks", what kind of feedback do you feel that you are owed?
Convincing LLVM to vectorize is still a problem in both languages. I do hope Mojo can make some headway there in the future. Especially since with MLIR they might be able to capture some higher level semantics Julia can't.
Did you read the Mojo code? It’s very messy and low-level dealing with explicit SIMD intrinsics and such.
Sometimes "showing the code" is not enough. Show me the benchmark.
The only other explanation is if you ran the non-simd Julia version under a single thread.
Without it, Julia starts single threaded, which means the code does all this work to enable multithreading and then doesn't get to benefit from it.
Would the C compiler automatically exploit vectorized instructions on the CPU, or loop/kernel fusion, etc? It’s unclear otherwise how it would be faster than Julia/Mojo code exploiting several hardware features.
IIUC, SIMD.jl only works because it only provides what is guaranteed by LLVM to work cross-platform, which is quite far from being able to use AVX2, for example.
I am keenly interested in Mojo, and have been running the local SDK for a few days on my Linux laptop.
I was also very keen on Julia a few years ago. I evaluated Julia as a true general purpose programming language: ML, DL, string processing, web use, etc., and it looked very good. Still, I didn’t switch from my go to languages Common Lisp, Python, and Scheme.
I have some hope, but limited expectations, that Mojo may become my one general purpose programming language. I started a template for a Mojo AI and general programming book. It helps me to write about new tech I am very interested in.
> Write Python or scale all the way down to the metal. Program the multitude of low-level AI hardware. No C++ or CUDA required.
I just thought they might already have something to show on that end...
has anyone seen programs like that?
Important detail that it runs on GPU.
sarcastic
[1] https://github.com/org-arl/InteractiveViz.jl [2] https://github.com/MakieOrg/Makie.jl
I expect Mojo's first release to be fast enough that it would get the Python folks using it over Julia and that the next Mojo release or future ones will get another 8x or 10x faster.
The same outcome happened with Bun 1.0 and already claimed the top spot in speed and compatibility in the node ecosystem in its first release and it isn't even done yet.
The least amount of effort to get something done much faster wins by default.
And then you’re stuck with a proprietary language “subscribe to download”? Not sure it’s the least amount of effort honestly.
By comparison, the simd-optimized Julia code (especially the first version) is significantly more elegant and transparent.
Impressively, the ComplexSIMD Julia class was defined in a few simple lines, from scratch. I wonder what the, apparently built-in, complex simd functionality in Mojo looks like under the hood.
> Why not develop Mojo in the open from the beginning?
> Mojo is a big project and has several architectural differences from previous languages. We believe a tight-knit group of engineers with a common vision can move faster than a community effort. This development approach is also well-established from other projects that are now open source (such as LLVM, Clang, Swift, MLIR, etc.).
https://docs.modular.com/mojo/faq.html#why-not-develop-mojo-...
that's cool but mojo literally just came out
There are always tradeoffs, and it usually takes a few weeks for people to come to terms with why Julia is unique.
Definitely falls into the fun category. =)
This might sound counterintuitive given that latency is a normal problem mentioned everywhere else about Julia. But, if you think about it, Julia compiled to native code a plot library from scratch in 15- seconds every time you imported it (before Julia 1.9 where native caching of code was introduced, and latency was cut down significantly).
This makes that problems where you would like to (for example) generate polynomials in runtime and evaluate then a billion times each, Julia can generate efficient code for ever polynomial, compile it and run it fast those billion times. C/C++/Fortran would have needed to write a (really fast) genetic function to evaluate polynomials, but this would have always (TM) been less efficient than code generated and optimised for them.
Edit: typos and added some remarks lacking originally
Only a few like Go ecosystem developers tended to take the time to refactor many useful core tools into clean parallelized versions in the native ecosystem, and to a lesser extent Julia devs seem to focus on similar goals due to the inherent ease of doing this correctly.
When one compares the complexity of a broadcast operator version of some function in Julia, and the amount of effort needed to achieve similar results in pure C/C++... the answer of where the efficiency gains arise should be self evident.
One could always embed a Julia programs inside a c wrapper if it makes you happier. =)
https://fortran-lang.discourse.group/t/fortran-is-faster-tha...
gcc is notoriously:
1. inefficient compared to the Intel or LLVM compiler
2. nondeterministic with -O3, which is why most people use g++ to check the code... and even then all bets are off on some hardware.
3. thrashes ram layouts, and slowly chokes to death if used as intended.
It comes down to the use-case, but fortran has killed too many to trust anywhere. =)
In general though it's just a question of which hoops you have to jump through for which language comparing C/C++/Julia/Fortran when using LLVM
The reality is that the Julia optimization was just a rewrite to use the same algorithm as Mojo, and that the Mojo code was heavily optimized.
One run, 7ms, 2ms? It is statistical fluctuation, not data, especially if it was run on "typical" developer laptop under "typical" session where browsers and other high-hitters are run in background and all these turbo-boosts and freq-governors are not turned off.
You need OS where almost all software (including most system services) are killed, CPU frequency is fixed (all power saving technology is turned off in both firmware and OS, I'm not sure it is possible on M-based Apple laptops, and many Intel-based laptops with castrated BIOSes are not suitable too).
You need warm-up loops to warm-up caches, or special code to flush caches, depends on what you want to measure.
You need to have tight-loop with your function called (and you must be sure, that it is not inlined by compiler into loop) which runs enough iterations to spend at least several seconds of wall time.
You need several such loop runs (10+, ideally), to have something which looks like statistics.
You need to calculate standard deviation and check that it is small enough (and you need to understand why it is not small enough if it is not).
Then it is benchmark.
Otherwise it is FUD.
It does not fix processes to CPU's, or set kernel governor to performance, and there are fluctuations from usage of the computer. But it does run the function for several seconds and returns the distribution of the runs (the little graphics underneath the benchmarks). It calculates standard deviation and if some runs are too small (sub-nano seconds) it emits warnings saying the results might be caused by inlining and constant propagation.
The differences in runtimes you refer to are from use of different machines or different routines, which is completely expected. They also argue they need to run the Mojo code in the same machine as the Julia code to be able to give meaningful results and comparisons.
While to someone outsider it might be seen as done without care, I can asure you that this people are used to take extreme care on how they do benchmarks. Again, it might just be that you're not familiar with the tooling developed to do it.
I do think there is more benchmarks needed to be done, as the Mojo code hasn't be optimised yet and none in that thread was able to run both the Julia code and Mojo code in the same machine (outside of the OP). But I'm sure this will be done (I guess rather sooner than later). :)
[1] Documentation of the package used for benchmark https://juliaci.github.io/BenchmarkTools.jl/stable/ Here you can find all the information you have said in your comment, and more, about reproducibility of benchmarks in different environments. White paper about the strategies used by the package https://arxiv.org/abs/1608.04295
> I don't know if there is more "compiler friendly" code with the same semantics for Mojo
The Mojo code here is from the official docs [1], so it's from the people best placed to know what the most "compiler friendly" code for Mojo would be, and what idioms they should use to get the best performance Mojo can provide.
The Julia macros @btime and the more verbose @benchmark are specially designed to benchmark code. They perform warm up iterations, then run hundreds of samples (ensuring there is no inlining) and output mean, median and std deviation.
This is all in evidence if you scroll down a bit, though I’m not sure what has been used to benchmark the Mojo code.
I used to be able to count on commenters understanding what was written even if they disagreed, but lately I see many comments confidently responding to something that wasn't relevant or even present.
One can say anything, but the Julia community takes performance and benchmarking really seriously.
These are not single runs of the code. The Julia code uses `btime` from BenchmarkTools, which runs many iterations of the code until a certain number of seconds or iterations is reached. The Mojo code uses `Benchmark` from a `benchmark` package, which I assume does similar things.
Beyond that, this is one person getting curious about how a newly released language compares to an existing language in a similar space, and others chiming in with their versions of the code. If you have a higher standards for benchmarks and think it will make a difference, you're welcome to contribute some perfect benchmarking results yourself.
They could have used @benchmark instead of the @btime macro, though. The first gives you the statistics, you asked for, whereas the second one is a thin wrapper around @benchmark, that just prints the minimal time across all runs.
Nevertheless the takeaway of this thread is pretty clear, even without @benchmark: The performance difference mainly stems from SIMD instructions.