[1] https://jochenschroeder.com/blog/articles/DSP_with_Python2/
For someone outside the field of DSP, please note that your post gives no context to the problem, nor to the input data. It's pretty hard to optimize something when we don't understand the structure of the problem or the input, and what exactly the code is trying to accomplish.
I remember taking a look at the post when it was posted some time ago, and having to give up pretty quickly because it was hard to make sense of it (in any language), and it seemed like I had to install specific Python libraries to (- hopefully, if everything works -) generate the input data. That takes it beyond "fun little optimization diversion to the day" to something that feels like work.
(Also, to anyone who hasn't read the post, I'll note that somebody did point to something obvious, as mentioned in the update to the post - a single suggestion that only needed superficial understanding of the code, that lead to a 3x speedup from OP's original Julia code.)
FWIW, someone with some domain knowledge on the Julia Discourse [1] was able to generate the input data, and posted code that they claim should be faster. They only mention the run time of the new code - 9 ms - so I don't know if, and by how much, it is faster in comparison. (Note that the code uses global variables at the end - which is fine there since they're using `@btime` and `$`s, but if using `@timed` instead, those should be enclosed in a `main` function or a `let` block.)
[1] https://discourse.julialang.org/t/this-post-claims-that-juli...
wxy[:,:,k] += mu*conj(err[i,k])*X
This should probably be wxy[:,:,k] .+= mu*conj(err[i,k]).*X
so allocations are avoided. This doubled the speed for me, although I don't what size inputs are realistic. The @benchmark and @profile macros are good for this stuff.I partially agree with this - that statement is true some of the time with some code, but in general there's a lot of nuance to that. (Though, I haven't seen that as a general claim much apart from in the initial "Why we created Julia" post describing what they were hoping to achieve; I more often see specific claims where people were able to port their scientific code into Julia that looks like Python and runs like C. But that may just be us having different bubbles online.)
My partial disagreement is from the fact that oftentimes, the "ease of Python" seems to get interpreted by Python coders as "as easy as Python is to me today" rather than "as easy as it was when I was a beginner" i.e. they expect to be able to port their Python knowledge directly somehow. Python has its own quirks and pitfalls, and over the first few months of people learn to work with/around them and write idiomatic Python code (and then it becomes second nature, and we often even forget that we had to learn to do this). Given a similar on-ramp, people can easily write Julia code that's at least comparable to C in terms of speed, even if it doesn't always match it.
But when you do need to match or exceed C, Julia starts asking for more in-depth knowledge of allocations, types and their performance characteristics, and the code starts deviating away from looking quite as readable as Python. Hopefully, you only need that in some deep internal functions or packed away in a library, but it _is_ a big caveat to that statement.
Try timing Julia's IO, or it's string processing, or it's hash tables, or its array allocation or it's GC - you know, nearly everything other that mathematical operations on floats or ints. Nearly everything leaves performance on the table.
Sure Julia is fast, but not on a C level in these areas. There is still lots of room for optimisation, though.
However, given this library seems to be about passing NumPy arrays to/from Julia, the application is almost certainly to expose numerical stuff from Julia to Python. Basically, the use case is similar to Numba, but it works by producing a C-extension from a Julia system image (if I understand correctly). The main advantage over something like Numba would be access to the Julia ecosystem.
In contrast, if your performance bottlenecks are IO, parsing and hash table operations, then Julia's performance will be in par with Python and get absolutely crushed by Rust (I have less experience with C, but would imagine Rust and C are about equally fast).
The major difference with Python is that Julia's implementation of these things _could_ be fast in a way that Python's can't. They just aren't.
Though I agree parsers haven't gotten much love in Julia. That said, this repo is saying it's for implementing NumPy extensions, and I don't think NumPy has many parsers it's using.
Anyway, I think it might be a good exercise for me to make a repo with a few test cases of where I think Julia could use some optimisations, and compare to Rust and Python.
fannkuch-redux, n-body, spectral-norm, reverse-complement: the fastest C wins by a large margin because it's an unreadable mess of vector intrinsics. If you really wanted to, you can do that in Julia too (and it might read better). Julia looks pretty even with C in the ones that were faithfully translated.