JNumPy: Writing high-performance C extensions for Python in minutes
github.com
github.com
Also why should I pick this over cython, pythran or numba? with those I don't need to learn how to optimise another language (and no just writing Julia code does not necessarily get you large speed ups and certainly not the speed of C in many cases).
I really wish the Julia community would stop overselling the language.
fannkuch-redux, n-body, spectral-norm, reverse-complement: the fastest C wins by a large margin because it's an unreadable mess of vector intrinsics. If you really wanted to, you can do that in Julia too (and it might read better). Julia looks pretty even with C in the ones that were faithfully translated.
Try timing Julia's IO, or it's string processing, or it's hash tables, or its array allocation or it's GC - you know, nearly everything other that mathematical operations on floats or ints. Nearly everything leaves performance on the table.
Sure Julia is fast, but not on a C level in these areas. There is still lots of room for optimisation, though.
However, given this library seems to be about passing NumPy arrays to/from Julia, the application is almost certainly to expose numerical stuff from Julia to Python. Basically, the use case is similar to Numba, but it works by producing a C-extension from a Julia system image (if I understand correctly). The main advantage over something like Numba would be access to the Julia ecosystem.
In contrast, if your performance bottlenecks are IO, parsing and hash table operations, then Julia's performance will be in par with Python and get absolutely crushed by Rust (I have less experience with C, but would imagine Rust and C are about equally fast).
The major difference with Python is that Julia's implementation of these things _could_ be fast in a way that Python's can't. They just aren't.
Anyway, I think it might be a good exercise for me to make a repo with a few test cases of where I think Julia could use some optimisations, and compare to Rust and Python.
Though I agree parsers haven't gotten much love in Julia. That said, this repo is saying it's for implementing NumPy extensions, and I don't think NumPy has many parsers it's using.
[1] https://jochenschroeder.com/blog/articles/DSP_with_Python2/
wxy[:,:,k] += mu*conj(err[i,k])*X
This should probably be wxy[:,:,k] .+= mu*conj(err[i,k]).*X
so allocations are avoided. This doubled the speed for me, although I don't what size inputs are realistic. The @benchmark and @profile macros are good for this stuff.For someone outside the field of DSP, please note that your post gives no context to the problem, nor to the input data. It's pretty hard to optimize something when we don't understand the structure of the problem or the input, and what exactly the code is trying to accomplish.
I remember taking a look at the post when it was posted some time ago, and having to give up pretty quickly because it was hard to make sense of it (in any language), and it seemed like I had to install specific Python libraries to (- hopefully, if everything works -) generate the input data. That takes it beyond "fun little optimization diversion to the day" to something that feels like work.
(Also, to anyone who hasn't read the post, I'll note that somebody did point to something obvious, as mentioned in the update to the post - a single suggestion that only needed superficial understanding of the code, that lead to a 3x speedup from OP's original Julia code.)
FWIW, someone with some domain knowledge on the Julia Discourse [1] was able to generate the input data, and posted code that they claim should be faster. They only mention the run time of the new code - 9 ms - so I don't know if, and by how much, it is faster in comparison. (Note that the code uses global variables at the end - which is fine there since they're using `@btime` and `$`s, but if using `@timed` instead, those should be enclosed in a `main` function or a `let` block.)
[1] https://discourse.julialang.org/t/this-post-claims-that-juli...
I partially agree with this - that statement is true some of the time with some code, but in general there's a lot of nuance to that. (Though, I haven't seen that as a general claim much apart from in the initial "Why we created Julia" post describing what they were hoping to achieve; I more often see specific claims where people were able to port their scientific code into Julia that looks like Python and runs like C. But that may just be us having different bubbles online.)
My partial disagreement is from the fact that oftentimes, the "ease of Python" seems to get interpreted by Python coders as "as easy as Python is to me today" rather than "as easy as it was when I was a beginner" i.e. they expect to be able to port their Python knowledge directly somehow. Python has its own quirks and pitfalls, and over the first few months of people learn to work with/around them and write idiomatic Python code (and then it becomes second nature, and we often even forget that we had to learn to do this). Given a similar on-ramp, people can easily write Julia code that's at least comparable to C in terms of speed, even if it doesn't always match it.
But when you do need to match or exceed C, Julia starts asking for more in-depth knowledge of allocations, types and their performance characteristics, and the code starts deviating away from looking quite as readable as Python. Hopefully, you only need that in some deep internal functions or packed away in a library, but it _is_ a big caveat to that statement.
The title is correct because it is a Julia numpy interface. Inside it seems is another project, TyPython, that provides an efficient Julia-Python-Numpy bridge, via C extensions.
One scenario would be that there is code already written in Julia, which you would like to expose to a broader audience. Although overall Julia's ecosystem is smaller than Python's, there are some niches within which it is highly developed. Library code accessible from the alternatives you listed is mainly limited to existing Python and C code. See the demo with `ParallelKMeans.jl`, which wraps a high performance Julia library performing K-means clustering, which could be a useful addition to the Python ecosystem.
Another scenario is that you want to write some custom code which makes use of the Julia ecosystem and then expose it to Python.
Quite the opposite according to my own experience.
Can you expand on that, so that it's possible to improve on them?
One can of course remove this in the compilation stage of PackageCompiler because PackageCompiler builds a new system image, and where BLAS is loaded is in the system image, so you can create a new from-scratch system image that is more lean. However, the tooling isn't quite there yet: right now the main way that's documented is something that extends the default system image, hence the large binaries. There's StaticCompiler.jl which does tree shaking so it makes small binaries, but it doesn't support most of the Julia runtime right now so it's limited in the codes it can handle. So right now the foundation all exists and it's at a usable state but definitely needs to improve.
Still, a large chunk of the memory being spent is not in the code that's produced or in the allocations happening in the code, but libraries that are loaded but not required. It's one of the downsides of focusing on interactivity first. The benchmarks in benchmarksgame don't reflect that, which I guess is up to interpretation/what's required - if the total amount of memory is a concern it's an important figure, if you only care about your core algorithm not having inherent allocation/memory problems, you probably won't care about the compiler chain as much.
For the love of god. Make it as fast as PHP.
My biggest pain point in computer science is that for every project I have to decide to either cope with PHP's butt-ugly namespace system or with Python's masochistic slowness.
Please, Python developers. Take a closer look at how PHP achieves its speed, hold all other development and copy whatever they do!
The result would be heaven.
That would be a step into the right direction.
6 more 30% improvements and it is on par with PHP.
Microsoft is misallocating its resources. The scientific ecosystem should be ported to .NET, with first class support for F#.
That is way easier to digest than the current ~5x performance penalty for using it.
I would probably not consider PHP anymore if Python was 50% of its speed.
I would never consider anything else than Python or PHP for web projects. Developer time is more precious than CPU time.
1. Maintain existing PHP codebase
2. Your team only knows PHP
You throw some PHP files on a server and have a working web application. As soon as you change a PHP file, the change is in effect on next reload of the page. You might be able to do the same in Python, but everybody in Python world is doing the long running thread approach, so you would be on uncharted territory.
That could actually work out nicely with the right libraries, just haven't yet heard of anyone doing that.
I think Python is inherently difficult to optimize because it allows some dynamic trickery that PHP does not. Additionally being large and complex does not help. You can't really pull an LuaJit equivalent. There sure are still many quick wins to be made for Python performance but not sure if it will ever compete with PHP/Lua/Julia and others on the performance front.
Not OP but I wish PHP was usable for scientific computing.
If I had better C/Rust/PHP internals insane macro skills I would integrate a dataframe library (Pola.rs?) into PHP to give us the basics for manipulating data.
Plenty more to add after that for matrix multiplication and the rest but I haven't thought that far ahead!
Good news: there are more than two programming languages!
But seriously, what situation are you in where your only choices are Python and PHP?
..which could be a heckuva lot friendlier and often safer than C!
The requisite python boilerplate is pretty ugly, though:
init_jl()
include_src('example.jl', __file__)
exec_julia('example.init()')
I also don't understand how jl_mat_mul gets mapped to mat_mul. jl_mat_mul = Pyfunc(jl_mat_mul)
Huh? from example.jl import mat_mul
and have it just work. (With the right libraries and import functionality of course.)jl_mat_mul = Pyfunc(mat_mul)
which makes a lot more sense.
I don't get it. Does this override python's numpy library with a julia native? If so does it implement all of numpy functionality? Or does this piggyback on normal numpy, redirecting accordingly, but also allows plugging extensions written in julia? The documentation doesn't shed any light at all. Neither do the demos.
Also, is this simply a python julia interface and the "high-performance" part of the description is just part of the usual juliaspeak? Or is there an actual "high-performance" component there that makes this more than a simple python julia interface?
Does anyone know, how JNumPy compares to PyJulia? Calling my Julia code via PyJulia seems to increase the startup time of Julia even more and it would be great, if JNumpy would do better here.
The announcement post [2] says "Calling Julia from Python, juliacall seems to have smaller overhead (in time) than julia". The benchmarks below that show the overhead on calling `identity` with `juliacall` being less than half of what it is with PyJulia.
[1] https://docs.juliahub.com/PythonCall/WdXsa/0.9.4/juliacall/ [2] https://discourse.julialang.org/t/ann-pythoncall-and-juliaca...
it's a breeze to use, header only.