Extending Python with Rust
maxwellrules.com
maxwellrules.com
If you test a single array operation with NumPy then it will compare favourably to a low level implementation, but where it generally will lose out is when you’re chaining together multiple operations which a compiler will vectorise more efficiently. For e.g. doing things like add two arrays then multiply them by another array and subtract a constant. If I put that in a C or Rust function I’d expect that to be auto-vectorized using fewer operations than what NumPy will do, and you of course also remove the overhead of the foreign function interface and control returning to the interpreter. On top of that, it’s usually trivial to drop in thread level parallelism in a lower level language, and NumPy doesn’t operate with core level parallelism out of the box for most operations, you have to use something like numexpr to achieve it.
I've personally used OpenMP + SIMD intrinsics or auto-vectorisation to great effect in scientific software and got performance >100x that of NumPy, but it's certainly not a free lunch. The question you have to ask yourself is normally “is it worth it”, and that’s something that can only be answered by profiling, looking at the overhead of maintaining the architecture, and understanding whether the increased friction with debugging and building is worth the hassle. Additionally, if you're doing linear algebra then generally if you're dropping to a lower level language you still will want to be using the same libraries that NumPy etc. uses unless you're exploiting some property of the matrix structure that can make operations more efficient in memory or compute. I worked on one codebase for e.g. where there was a block matrix structure and a hand written implementation of a solver was written because much more efficient operations could be performed with knowledge of that structure.
The main bottleneck that C-extensions solved for us then was marshaling. Instead of doing something like Numpy --> custom algorithm --> Numpy --> custom algorithm in Python, we moved it all to C. This way, we decreased the number of times we had to marshal data & we reduced the amount of data that needed to be marshaled.
IIRC we gained 10-100x improved performance in all scenarios where we used C-extensions. Before doing so we tried to use Numba, as it seemed to offer a "free lunch", but we never got it to work.
The final cherry on top that let us surpass real-time speed was replacing a graph algorithm with another that had a better time complexity for larger graphs. Before doing that switch, we'd still be faster than real-time for _most_ inputs but would lag behind severely whenever we had a dense input graph come in.
The service ingests data from a timespan T spanning M minutes. Before the next timespan T+1 is ingested, the service has to process the data from timespan T in <M minutes otherwise we fall behind. What I mean with "faster than real-time" is that it has to process the data faster than it is ingested.
Op is better off recompiling numpy to target his specific hardware thank trying to speed it up using rust. This would be to try different underlying math implementations that are linked with numpy ( blas, mkl, eigen). Each of these haha several internal simulation frameworks to optimize not just the actual instructions but memory layout for various math kernels.
I'd recommend using Numba first, though.
* Numba: https://numba.pydata.org/
* JAX: https://jax.readthedocs.io/en/latest/notebooks/quickstart.ht...
E.g. I’d expect for small/medium then re-writing as a fold over the input arrays would be fastest because you’d only traverse once & everything would fit in cache.
However if the 3 arrays combined size is larger than L1 cache, I’d be strongly tempted to bet on the naïve 3 operation approach to be faster. You’d save so much time on cache line flush / reload activities.
But then if 2 arrays are larger than L3 i’d expect going back to the fold / single traversal to win again because the iterating 2 arrays at a time behaviour is no longer any different to iterating 3 at a time.
Untested hypothesis.
Including another for would spell out to "for for example"
I think the bigger issue is remaining cache-friendly in the face of multiple operations. If you do the operations one at a time, you’re doing multiple passes over the data.
In any case, the approach that Polars[0] took seems like a pretty good solution, if difficult to implement well.
Polars is a dataframe library, and has a “lazy” API. This builds up a graph of operations, runs it through a query optimizer, and then executes them all at once. This allows for parallelization, optimizing expressions, and fewer passes over the data. The downside is that it’s nontrivial to implement a query optimizer.
[0]: https://pola.rs/
Thankfully there are GitHub Actions and similar tools that help with this.
Cython extensions on the other hand work fine with the MinGW GCC compiler provided by Anaconda, although it's pretty old (version 5).
If you're so inclined to revisit this using mingw with rust you can see how here https://rust-lang.github.io/rustup/installation/windows.html
This is far easier to automate imo
If I remember correctly I tried installing the MinGW based Rust toolchain too, but I'll try again with the instructions you've given me.
And with a lack of ARM runners by default with GH actions, you'll most likely be paying for your own CI instances (or wait forever for cibuildwheel cross compilation/qemu). Also for others doing this, the rust-cache GH action saves a lot of rebuild time too.
Python has neither standard tooling (de facto nor de jure), has no consensus within community where everyone seems to invent their own yet-another-dependency-management-tool and no culture of preserving strong backwards compatibility (never mind the core 2->3 transition).
But we'll agree python's dep management is leagues ahead of javascript and R, though. Right?
docs: https://docs.greptime.com/user-guide/coprocessor-and-scripti...
code: https://github.com/GreptimeTeam/greptimedb/tree/develop/src/...
CREATE FUNCTION pymax (a integer, b integer)
RETURNS integer
AS $$
if a > b:
return a
return b
$$ LANGUAGE plpython3u;
https://www.postgresql.org/docs/current/plpython.htmlhttps://ramanlabs.in/static/blog/Generate_Python_extensions_...
Do you mean maturin? Rumpy is TFA's demo project.
The surface plot brings my computer to a crawl. And, I'm not sure how to make it interactive. I gave up on it a few days ago.
Could have tried a Rust module, but ended up throwing together a custom Rust plotter in a few hours, with translated Python code. (I already have a basic WGPU-based rendering engine, with EGUI for UI). It runs much faster. They're both imperitive langs, so you can copy and paste, replacing `**` with `pow()`, `np.exp(x)` with `x.exp()` etc. And indexing with x[i][j][k] instead of vectorized numpy ops.
Ie the program compiles and runs in release mode faster than Python/numpy could perform the calcs. And the 3D graphics are smooth instead of a slideshow. And, I can make it real-time interactive since I'm using a low-level lang that integrates with GPU APIs directly. I'm not sure how feasible that would be in Python.
Note that there's currently no docs or template/examples, and I'm rapidly breaking the API.
The surface plots are just meshes of a grid divided into triangles. When I need to manipulate them, I re-gen the meshes from a nested 3D array.
So, not a plotting lib at all; a flexible 3d and UI lib. Could probably made usable by others with a basic example of how you interact with the render and UI.
Have you tried the Plotter lib? I haven't used it, but looks nice from a skim. https://plotters-rs.github.io/book/basic/draw_3d_plots.html
I've come across Plotters before but last time I checked, the documentation was pretty sparse and I didn't really relish the thought of trying to parse though and tweak the examples at that time. It looks like they're gradually filling out the docs now, though, so maybe I'll give it another go.
However, it won't solve your problem of Python not being fast enough doing the calculations.
I think the big issue with Julia here is (other than the improving(?) JIT slowness (IIRC it was faster to compile and run a Rust program than JIT a Julia script), is I'm not sure how I'd interface with the GPU for graphics, compute shaders, and a GUI.
Julia's mathematical syntax is best-in-class; I wish more langs used something like that.
As for JIT, just today there was a PR that was merged that makes Julia cache and reuse binaries of packages (https://github.com/JuliaLang/julia/pull/47184). It won't be out until the next release of Julia, but it's a pretty major improvement to not JIT fully inferred package calls.
TL;DR:
* If you're wrapping existing C library, I'd use Cython.
* If you're wrapping existing C++ library, I'd use PyBind11 (no personal experience, but it's based on Boost::Python, which I have happily used). Cython in theory does C++ but it's a frustrating, limited experience.
* If you're writing a tiny library and you don't know Rust, and you're not worried about memory safety, Cython is nice.
* For anything involving writing extensive new low-level code, Rust with PyO3. Memory safety _will_ bite you in the ass. Concurrency is vastly easier with Rust. You get a package manager for dependencies. You're not writing a pile of code in a language without good tooling (Cython).
Long form, with more alternatives and use cases: https://pythonspeed.com/articles/rust-cython-python-extensio...
https://github.com/numpy/numpy/blob/22e683d84f2584a6f9a57b2c... + https://github.com/numpy/numpy/blob/22e683d84f2584a6f9a57b2c...
To know what they do, you need the source for INPLACE_GIVE_UP_IF_NEEDED (https://github.com/numpy/numpy/blob/b222eb66c79b8eccba39f46f...) and PyArray_GenericInplaceBinaryFunction (I don't even know where that's coming from, it's not defined in Numpy, maybe it's part of the Python interface?).
In the end, both are unreadable in their own way. I personally prefer the Rust version above to the macrofied C version that's in Numpy but that's a matter of taste. I'd also trust the safe Rust implementations more than the C implementations because of the memory management guarantees Rust provides, though I suppose for simple operations like multiplication it'll be easier to make the program safe enough in C.
https://github.com/martinxyz/progenitor/blob/85260/crates/pr...
I suppose what is meant is faster (also follows from the diagram?). But it is still not a dramatic gain for many use cases. This shows how non-trivial the python performance calculus: pure python, versus numpy python, versus compiled c/c++ or rust. People who want to speed up python should really look whether numpy helps before complicating their codebase more.
But there are more benefits to those bindings besides performance so its really nice to see the expanding options
A pure rust implementation will likely always be slower by virtue of not using the same tightly designed optimized code.
Side note: A fun implementation detail of numpy is that after you install it from pypi, it does a user side compile of some of the modules on first import. Which means you need to be somewhat careful if you ever relocate an install of it to a new machine
Deprecated since version 1.20: The native libraries on macOS, provided by Accelerate, are not fit for use in NumPy since they have bugs that cause wrong output under easily reproducible conditions. If the vendor fixes those bugs, the library could be reinstated, but until then users compiling for themselves should use another linear algebra library or use the built-in (but slower) default, see the next section.
Source: https://numpy.org/doc/stable/user/building.htmlhttps://numpy.org/doc/stable/release/1.21.0-notes.html
> With the release of macOS 11.3, several different issues that numpy was encountering when using Accelerate Framework’s implementation of BLAS and LAPACK should be resolved.
I know nothing about Rust, but I’ve noticed that there’s a top post on HN every day that followed this format.
So it's a bit of a meme, sure, but there is a good reason for it :)
However another genre is more interesting, to me at least - utilising existing interfaces to extend software people already use without having to make them throw out what currently works and take the risk of replacing it with a recently implemented version. In this case someone implemented a Python module in Rust, in others entire Linux device drivers have been implemented in Rust. Lower-level programming in embedded systems seems like a good application too, but I don't know how many architectures Rust can target and is officially supported on.
So to summarise, the two genres[0] I've identified are, roughly:
- I did a RIIR[1] of 50% of grep's functionality for fun
- Here is how you can accomplish a common task in C using Rust instead
As I said, not a Rust dev and frankly I'm quite intimidated by all the rants people had about fighting with "The Borrow Checker" which sounds like a ferocious Elden Ring boss. But I am tempted by the way it allows safe (or safer) software to be written.
[0] - both are valid, and more exist but these are two common ones that came up
[1] - RIIR = "Rewrite It In Rust", which was a bit of a meme for a while as a bunch of devs got excited about Rust and launched projects of varying completeness reimplementing things.
In my experience, those "fighting" the borrow checker are often novices misunderstanding the language model or lacking the experience necessary to effectively use it, or extremely advanced programmers running into compiler limitations and bugs. In most Rust code there's very little fighting the borrow checker.
Compare it to people saying they're fighting the compiler or fighting the interpreter when they try to multiply a string by an array and only errors come out, or trying to reassign a const value and cursing at the compiler for getting in their way. The error messages are often unhelpful and unclear (though I have to give Clang and Rust that their error messages are quite good in these days) but the core "fight" is trying to do something that doesn't make sense within the context of the programming language.
For Rust, there are additional limitations on top of what your average JS/C#/Java/Python programmer will be used to. These limitations are often also present in C (you can't just share memory between threads willy-nilly, or pass around freed pointers!) where they're classified as undefined behaviour, hopefully with a warning in the console, but in Rust you must deal with the error that's been detected.
The borrow checker is an extra constraint that's present in many modern C++ code bases as well, though the compiler lacks proper analysis support in many areas. It takes some getting used to memory ownership and the limitations and possibilities associated with it, but in many cases listening to and understanding the borrow checker will give you much better code (more correct code but also often clearer code) than ignoring the warnings like you would in C++.
For most programming languages, a beginner needs to learn 1) installing/calling the tooling 2) the core language syntax 3) the special magic features of the language and 4) how to distribute the compiled code. With Rust there's an intermediate step between 2 and 3, the borrow checker, which beginners tend to underestimate or dismiss (I certainly did) despite the warnings in any guide. It's not especially complicated, but it's an extra step other language might lack.
I recommend giving learning Rust a go. Don't be like me, don't skip the uninteresting parts of the Rust Book (https://rust-book.cs.brown.edu/, chapter 4 is what I'm referring to), you'll only find yourself getting more frustrated. Chapter 17 (fearless conspiracy) is where I fell in love with the language, but to get through it I had to go back and actually read about ownership.
That sounds like a very different kind of book.
OT: I quite like the concept of the borrow checker saving me from shooting myself in the foot, it's just a shame that Rust has traits, without those it might have been a nice language to use.
users.remove(expired_user);
Which gives me flashbacks to Java and makes me want to saw open my skull and pour bleach over my brains. users is data, it should not contain logic.well yeah as a n00b I would be both of those
https://github.com/nushell/nushell/blob/768ff47d28ec587b689a...
And comments like yours are following one as well.