Faster Python calculations with Numba
pythonspeed.com
pythonspeed.com
I've been using it in a python graph library to write graph traversal routines and it's done me very well: https://github.com/VHRanger/nodevectors
The best part is the native openMP support on for loops IMO. Makes parallelism in data work very efficient compared to python alternatives that use processes (instead of threads)
And I agree that it's not actually usable everywhere, since the support for numpy's feature set is actually quite limited, especially around multidimensional arrays. I had to effectively rewrite my logic to make use of numba. Still it is pretty worth it imo, given how it can add parallelism for free. And conforming to numbas allowed subset of numpy usually results in simpler and more efficient code. In my case I ended up having to work around the lack of support for multidimensional arrays but ended up with a more efficient solution relying on low dimensional arrays being broadcasted, reducing a lot of duplicate computations
"Fits very few use cases" LOL okay without numba there's no UMAP and HDBScan and those are pretty popular and important libraries that come to mind just off the top of my head...
Also, claiming Cython is well documented also gets a huge LOL from me as someone whose actually written a bit of Cython.
I've succesfully deployed numba code in an AWS lambda for instance -- llvmlite takes a lot of your 250mb package budget, but once the lambda is "warm" the jit lag isn't an issue.
That said, if you absolutely want AOT you'll have to use Cython or some horrible hack dumping the compiled function binary.
I like numba, but cython is clearly used more in the popular packages
In the past, that was also possible for AOT compilation [1], but that technique broke during some update and it seems like there is no one left who knows how to fix this.
If you want to specify the type (for example for aot or just because you want to make it clear) then the call signature is less flexible.
In short, pick any random Python library, you’d find there are very few places you can jit accelerate something effectively. It is for numeric.
Even for numerical code, it is more like writing C functions than say C++ (with classes etc).
But it does makes accelerating vectorized code very easy. Even if you have a function that uses Numpy, it is likely you can speed it up using Numba with a decorator only.
But when it doesn’t work, it might often be not very clear why you can’t until you get some experience.
Though if any numba developers come across this, I’d advise them to plan their upgrade process a little more carefully. The current numba conversion for python lists is due to be depreciated in favor of typed lists. But the typed list (List()) method is still too buggy. I experienced huge delays when popping and appending elements. Please flesh out these new methods before adding a bunch of warning messages about depreciation to the current builds.
numba lets you inline with a subset of python which is then compiled with llvm producing something very similar to what you would get if you applied a bunch of regexes to that subset of python to convert it from python to C. (with special bindings for numpy arrays, since they have special importance in these domains)
numba is specifically targeted at things like core numerical algorithms that are typically coded in C and fortran, and are typically comprised of solely for loops and basic arithmetic. JAX is more targeted at high level machine learning applications where the end user is stringing together more high level numerical algorithms.
i suspect that JAX would be a bad fit for custom computer vision or numerical algorithms that are used outside of the use-case of doing neural networks work.
I am currently doing some tests introducing JAX in a large numerical code base (that was previously using C++ extensions), we are not using autograd nor any deep learning specific functionalities. Having seen actual numbers, I can tell you that JAX on CPU is competitive with C++ but produces more readable code with the added benefit of also running on GPU. However, it does introduces some constraints (array sizes cannot be too dynamic) so, if you are not also planning on also targeting GPU, I would probably focus on numba.
Jax is a great general purpose numerical computing library.
also a little bit surprising to see how immature and fragmented the python gpu numerical computing ecosystem is. everybody bags on matlab, but it has been automatically shipping relevant operations over to available gpus for years.
In Numba it only recompiles for dimension changes i.e. if shape changes from (7, 10) to (2, 10, 12). However numba does not integrate AD so if you need that, JAX is probably your best bet unless Enzyme matures and you want to look into integratig it with numba.
I've done a pretty extensive amount of development in Cython. I agree, it's very nice to be able to incrementally type Python code but I've found that using the language to its full potential requires a good understanding of C programming, especially for debugging.
[1] https://pythonspeed.com/articles/rust-cython-python-extensio...
I don't think that they even compete or solve the same use case tbh. Not sure why they're being constantly compared to each other.
We build a very computationally intensive portion of our application with it and it has been running in production, stable, for several years now.
It’s not suitable for all use cases. But I highly highly recommend it if you need to do somewhat complex calculations iterating over numpy arrays for which standard numpy or scipy functions don’t exist. Even then, often we were surprised that we could speed up some of those calculations by placing them inside numba.
Edit: ex of a very small function I wrote with numba that speeds up an existing numpy function (note - written years ago and numba has undergone quite some amount of changes since!): https://github.com/grej/pure_numba_alias_sampling
The ELI5 is that numba allows you designate specific functions in your application for compilation. That tells the library to run a just in time compiler on it. If you can do that successfully, you can wind up with that portion of your code base running at or near C-speed. There are a lot of gotchas and limitations, and it only works for a subset of the python language. But when it works, it can make loops over numpy arrays very fast, when they would normally be extremely slow in pure python. In a lot of cases, just a few compute bound processes are holding back the whole application. Numba allows you to improve those portions in a very targeted way, if you have a suitable use case.
Because my code is parallelized, I had to figure out how to trigger the compiler before forking because then each worker process had to do the compilation itself. There’s some decorator syntax to tell numba to precompile for specific numeric types, but I found it easier to just call the functions when you import them.
All told, it’s way easier than it would have been to implement this hotspot in C++ or something.
Suppose we have two very large vectors and we want to take the dot product of them. If we store these vectors in a standard Python list, write pure Python, and try to run it in CPython, it is very, very slow, at least relative to how fast almost any other non-dynamic scripting language could do it. (I have to add all those caveats because this is easy mode for a JIT and PyPy might blast through this. But CPython will be very slow.)
Suppose instead we have a data type that is a C array under the hood, and all we do is this:
array1 = load_from_file("something")
array2 = load_from_file("other_file")
dot_product = array1.dot_product(array2)
Suppose "load_from_file" loads into an optimized data structure that provides Python access, but under the hood is stored efficiently, and the "dot_product" method leaves Python and directly performs the dot product in some more efficient language.In this situation, you're running a bare handful of slow Python opcodes, a few dozen at the most, and doing vast quantities of numerical computation at whatever the speed of the underlying language is, which in this case can easily use SIMD or whatever other acceleration technologies are available, even though Python has no idea what those are.
That is the core technique used to make Python go fast. There are several ways that things get from Python to that fast code and underlying structure. The simplest and the oldest is to simply do what I laid out above directly, with external, non-Python implementations of all the really fast stuff. I believe this is the core of NumPy, though I'm not sure, but it's certainly been the core of a lot of other things in Python for decades.
There are also approaches involving writing in a Python-like language, but one that can be compiled down (Cython), approaches around trying to compile Python directly with a JIT into something that under the hood is no longer using CPython's interpretation structure (PyPy), and other interesting hybrid approaches.
CPython is generally a slow language, but it isn't always a big deal in practice because if you write Python code where a reasonable amount of most lines you write go out to faster code, the slowness of Python itself is no big deal. For example, you can write decent, performant image manipulation code in Python as long as you're stringing together image manipulation commands to some underlying library, because in that case, for each "expensive" Python opcode, you're doing a lot of "real work" at max speed. I scare-quote "expensive" because in that case, the Python opcodes may be a vanishing percentage of the cost of the work. The problem with numerical manipulation in pure CPython is that it swings the other way; each Python opcode requires many assembly-level instructions to implement, but it's inefficient to run all those instructions just to add two integers together, a single assembly opcode.
Arguably the dominant factor in sitting down with any programming language and writing reasonably perfomant code on your first pass is understanding the relative costs of things and just being careful not to write code that unnecessarily amplifies the costs of doing the "real work" with too much bookkeeping. The same 1000 cycle's worth of "bookkeeping" can be anything from a total irrelevancy if your "payload" for that bookkeeping is many millions of cycles worth of work, to a crippling amplified slowdown if your payload is only one cycle's worth of work. Even a "fast" language like C++ can be brought to a crawl if you manage to use enough abstractions and bounce through enough non-inlinable method calls and through the requisite creation and teardown of function frames, etc. just to add two numbers together, and then do it all again a few million times.
https://github.com/synapticarbors/ndarray_comparison/blob/ma...
Most surprising among them was how fast pythran was with little more effort than is required of numba (still required an aot compilation step with a setup.py, but minimal changes in the code). All of the usual caveats should be applied to a simple benchmark like this.
https://docs.rs/inline-python/latest/inline_python/
use inline_python::python;
let who = "world";
let n = 5;
python! {
for i in range('n):
print(i, "Hello", 'who)
print("Goodbye")
}
The nice thing about this, if you get the power of Python to handle your data loading, scrubbing, validation, etc and then spend the bulk of your high performance work in Rust.Layout is king, so structuring your code to have compact, native layout will not only make it much easier to call compiled code, it will vastly reduce your memory requirements.
import cffi
ffi = cffi.FFI()
chunk = ffi.new(f'int[{5_000_000}]')
chunk can now be used as a fixed size array. And chunk can be passed by reference in native code (Zig, Rust, Fortran, etc).Don't get me wrong I like rust, but I would always write the high performance code to be a module that can be imported in python, write everything in rust and the use python for some specific functions only.
Contrast that with Lua where embedding and extending are about the same difficulty, there one is more likely to pick the technique that best fits the situation.
One thing that is really nice about calling into Python when you need it but otherwise using your calling language and runtime is you get all the semantics of your calling language to structure the codebase and use Python as a library where needed. This is esp nice when you have security or concurrency properties that you need to hold, or if you were bootstrapping a project with Python and all of its wonderful libraries but eventually will replace or reduce the script code.
One semantic nit, you probably do understand it, you just don't like it. If I have a big ETL that has lots of parts that need to run in Rust, but the cleanup code would be much easier to write in Python, I don't need a to turn my ETL into a library.
Julia was also faster in my experiments.
I have tried Julia but I found the promised "everything will be as easy to write as python but as fast as C" is often not true. There are quite a few pitfalls (e.g. don't use abstract data types), which one has to consider and code can start looking looking more like cython than python if one wants to optimise for speed. I also dislike some of the design decisions they made (in particular making multiply a matrix operation and using the '.' for broadcasting which is incredibly easy to overlook and one of the main sources of bugs in matlab code in my experience).
That said there are great things happening in the Julia ecosystem, just don't promote it as the be all of scientific computing all the time.
And when thinking about something more than optimizing a single function, but about a large codebase and the context in which you need it, Julia is rarely a compelling option.
I did a comparison to cython and Julia here: https://jochenschroeder.com/blog/articles/DSP_with_Python2/
Did you have a look at JAX? It would be interesting to see how well it would compare.
The exercise for me was to see how I fast I could get the julia code with my background, I am quite honest about the fact that I have less experience in julia, however for me the whole reason to look for something else then cython is that it takes so much background knowledge to get the ultimate speed, if I have to do the same for julia I don't really gain anything. I would argue the amount of optimisations in the julia code are already more than what I had to do in pythran.
Numba just worked, with a very straightforward implementation of my calculation as well.
Really cool piece of kit.
I think one should just vectorize the solution from the get-go! np.accumulate looks like the right way to go about things in this example.
Btw the J implementation would be >. /\ and the k implementation would be |\ which is I think much better than the for loop anyway.