PyO3: Rust Bindings for the Python Interpreter
github.com
github.com
It just calls the Rust's chrono library that does the parsing and wraps the result in a Python object. You can do it for any Rust library, it's very, very easy!
The only slightly complicated part is the distribution. You need to use https://github.com/PyO3/maturin or https://github.com/PyO3/setuptools-rust, and of course, you need to have Rust installed on the wheel-building machine.
Feel free to use this repo as a reference if you want to build a similar thing. The code is commented, and there's a working GitHub action that builds the wheels for all platforms and uploads them to PyPi: https://github.com/gukoff/dtparse/tree/master/.github/workfl...
Because our problem was just about reformatting we ended up reading the CSVs in binary mode and using struct to extract the relevant values from the date time fields. But if we needed to do actual date logic something like this would perhaps be useful (but there other fast date time libraries out there, I've been a fan of pendulum for some tasks).
If you have a batch SLA of 1 hour, and your currently spending 50-70 mins to complete the batch and 20 minutes of that time is spent date parsing and you can reduce it to 5 minutes that's an big win.
These JSONs represented the flight information. They included multiple datetimes, such as the scheduled departure/arrival time and the real departure/arrival time of a flight.
The first bottleneck was JSON deserializarion/serializarion. At that time we solved it with ujson, and now there's the even more performant orjson.
The second bottleneck happened to be datetime deserializarion. And we solved it with ciso8601 - luckily, these datetimes were in ISO8601. But this bottleneck later repeatedly occured in the other components and became an inspiration to write dtparse :)
I don't remember the numbers, but caching + using ciso8601 was essential to manage the peak load (maybe 50k trades per sec ?).
I was looking at PyO3 a few months ago, after discovering the orjson python (with rust inside) library and radically speeding up an auto-ML app for work.
I really enjoyed starting to learn Rust, but found the process to embed in Python to be rather intimidating. Looking forward to using your repo as a reference, and love the dtparse work you've done.
to be fair ciso8601 only parses iso8601 datetimes, but that's enough for 90%+ of my use cases.
On my machine, ciso8601 always runs in 240ns, and the Rust lib median time is 1250ns.
You can run a benchcmark too! Just call pytest, and it will generate an .svg report: https://github.com/gukoff/dtparse/blob/master/tests/test_per... (you'll need to pip install ciso8601 pytest pytest-benchmark[histogram])
I ended up looking at a bunch of different ways of processing timestamps in Python: strptime(), string parsing, regex, datetime.isoformat(), NumPy, Pandas, and more. I got a 46x speedup using datetime.isoformat(). Other approaches got anywhere from 4x to 40x speedup, and a couple approaches were an order of magnitude slower than strptime().
My takeaway was there's no substitute for profiling the actual code you're running, and focusing on the specific bottlenecks in your own project. I wrote this up in a blog post if anyone's interested, "What's faster than strptime()?"
Major caveat is timezone handling, but this only applies in a subset of situations
Now I'm starting to use it as part of the Python memory profiler I'm working on (https://pythonspeed.com/fil), in this case to call in to the low-level Python C API which PyO3 includes bindings for in addition to its high-level API. This kind of usage is more like writing C, except with the benefit of having high-level APIs (for GIL holding, but also object conversion) available when I need it.
So basically you get safe, high-level, easy-to-use APIs, with fallback to low-level unsafe APIs if you need them.
Highly recommend trying it out.
In general I suspect it's the usual "NumPy arrays are fast, everything else you better be getting a sufficiently large boost from the low-level code to justify conversion".
For the thing I prototyped in Rust, it was wrapping the `ahocorasick` crate which was in fact faster than `pyahocorasick` which is written in C or Cython or something. Both have similar conversion costs, probably, so it came down to "for lots of data the Rust version was faster".
Or just be sure to enable the DFA option if you can afford it. It looks like the Python library is just the standard NFA algorithm.
Next step is trying alternative approach, but if that alternative doesn't work I'm going to see about wrapping your package for Python.
Thanks for all your work on it!
E.g. `df["column_name"].values` will you get you a NumPy array.
There’s a setuptools rust [0] extension package that can be used to hook the compilation of the rust into the wheel building or install from source. Maturin [1] seems to be regarded as the new and improved solution for this, but I found that it’s angled toward the using python from rust.
There’s also the rust numpy [2] package by the same org which is fantastic in that it lets you pass a numpy matrix to a native method written in rust and convert it to the rust equivalent data structure, perform whatever transformation you want (in parallel using rayon [3]), and return the array. When building for release, I was seeing speed ups of 100x over numpy on the most matrix mathable function imaginable, and numpy is no joke.
I think there is a lot of potential for these two ecosystems together. If there’s not a python package for something, there’s probably a rust crate.
If anyone is interested the python package that I'm building with some rust backend, its called pyrogis [4] for making custom image manipulations through numpy arrays.
[0] https://github.com/PyO3/setuptools-rust
[1] https://github.com/PyO3/maturin
[2] https://github.com/PyO3/rust-numpy
> There’s a setuptools rust [0] extension package that can be used to hook the compilation of the rust into the wheel building or install from source. Maturin [1] seems to be regarded as the new and improved solution for this, but I found that it’s angled toward the using python from rust.
> There’s also the rust numpy [2] package by the same org which is fantastic in that it lets you pass a numpy matrix to a native method written in rust and convert it to the rust equivalent data structure, perform whatever transformation you want (in parallel using rayon [3]), and return the array. When building for release, I was seeing speed ups of 100x over numpy on the most matrix mathable function imaginable, and numpy is no joke.
What sort of algorithm was that? Generally getting 100x speedup on vectorized code is highly unusual even using handcoded c++. So I suspect it was quite loop heavy? In those cases I have also seen very significant speed ups.
I have been using pythran [1] for speeding up my python code. It generally achieves extremely good performance. I have blogged about it here [2] and recently a member used pythran to speed up some nbody benchmarks [3] which was used in an article to argue for using compiled languages.
That said I find pyO3 quite exciting and have been contemplating to try it with some of my projects. [1] https://github.com/serge-sans-paille/pythran [2] https://jochenschroeder.com/blog/articles/DSP_with_Python2/ [3] https://github.com/paugier/nbabel
I’m checking out that post later, I’m trying to make my package easy to build on, so being able to write extensions with Pythran would be another great option for speed ups. Thanks
def threshold_pixel(img, thr): out = np.zeros_like(img) o = np.mean(img, axis=-1) out[o>thr] = 255 return out
This runs in ~30ms for a (1024,1024,3) array using numpy on my machine. Using pythran (note I had to explicitely write out the loop for out[o>thr] =255, due to a bug, that I found and just reported), I get a speed of 6.ms (with openmp) and 9ms without (I did not tune the openmp, but this should yield a much higher speedup).
P.S.: Just had a look at your project, very cool, I have to try that
Thanks Py03 team :)
Learning PyO3, I cobbled together a sample project[0] to demonstrate how some functionality works. It's a little outdated (uses PyO3 0.11.0 compared with the current 0.13.1) and doesn't show everything, but I think it's reasonably clear.
One thing I noticed is that passing very large data from Rust and into Python's memory space is a bit of a challenge. I haven't quite grokked who owns what when and how memory gets correctly dropped, but I think the issues I've had are with the amount of RAM used at any moment and not with any memory leaks.
Write code in python and transpile to another language (could be rust) and then import it back into python
https://github.com/adsharma/py2many/tree/main/tests/expected
Figuring out a mapping between a subset of a compiled language and a subset of statically typed python should be possible.
The hard part is mapping standard library. I suspect something like nim might have an advantage there.
Compile your Rust code to wasm to circumvent having to compile for different architectures.
It lives up to that claim. (I had issues with return object typing when going between Python/Rust at first but those are more consistent now)
[0] https://github.com/bytecodealliance/wasmtime/issues/2273
[1] https://github.com/bytecodealliance/wasmtime/issues/2274
O
O = Py < |
O
or Py (vi) O
||
O = Py
||
O
or Py (ii) O
Py < > O
O
heh!Has anybody used this in conjunction with a python framework? Django, fastapi or something?
> often these are a struggle with scaling interpreted languages compared to other lower level languages
Not sure what's meant by this
I'll give it a shot tonight and see how it goes. Now I'm curious.
Works really nicely, although given how little work I'm doing in the Python side I honestly prefer using Rocket instead of FastAPI and then using pyo3 to call the Python library in Rust, rather than the other way around.
I'm new to rust, but I'll check out Rocket. Cheers