Fixing Python Performance with Rust
blog.sentry.io
blog.sentry.io
The reason is that it can automatically generate C API code without FFI, because FFI calls are slower (https://gist.github.com/brentp/7e173302952b210aeaf3) so there is less overhead. You obviously care about overhead here.
Nim's pymod is already a python module and you can send strings and numpy arrays to Nim for fast processing.
I wish you could send bytes from python3, but that's not implemented yet.
In general, you can get to less overhead with Cython right now because it's a much more mature project and because you can be more flexible with defining the point where you drop from python to a faster alternative. I wouldn't use that Nim/Python lib in production unless you have time to help with its development. It's good for personal projects though.
Debugging and writing cython that is fast is feally not a very pleasant experience and the tooling is not great. Let alone that a real ecosystem exists. There is not even a good way to deal with dependencies at compile time.
> If a language's lexical structure is so tortured that I need a special tool to grep it, I'm going to skip it and go back to the sanity of C++. [2]
Nope, nope, no way.
[1] http://nim-lang.org/docs/manual.html#lexical-analysis-identi... [2] https://news.ycombinator.com/item?id=8936542
That way you can mirror, say, C style while writing the bindings either manually or automatically and use them as if they were normal nim elsewhere.
If I'm reading it correctly they're just using CFFI to load and call a shared library - it's really not embedding anything. The fact that the library was written in Rust is interesting, but as far as Python is concerned it could easily have been written in any language that can create a shared library.
There are examples of each of these compared to one or the other but it would be nice to see all of them compared in a few benchmarks.
Of which there are not that many that would make a shared library that can be safely loaded into a Python process. Traditionally this was limited to C and C++. As far as embedding goes: Rust like most things needs minimal runtime support and that is embedded in the dylib.
Here's how to do it Haskell, for example: https://downloads.haskell.org/~ghc/7.6.3/docs/html/users_gui...
That said, if optimization gets to the point where re-writing in a different language is a good idea, most people just jump to C or C++, or in this case Rust, because they're the fastest.
If they are running multiple Docker containers, the improved efficiency would allow them to run more containers per host computer. Maybe that's what they meant, although it's not really horizontal scaling either.
Performance has nothing to do with scalability whatsoever, by the way.
If we can do more with less machines it's no different than doing more with more machines. We added triple our infrastructure to handle the Python load while we resolved this CPU issue, and that wasn't "scaling" (literally, it causes other scale concerns).
Fixing the root cause let's us drop all of those new machines as well as some older ones. The scalability of the system has greatly increased because of this and other factors -- primarily that many or all systems aren't actually "horizontally" scaleable.
Scaling Horizontally - spending resources on adding more nodes to a cluster, so more work can get done in parallel. Scaling Vertically - spending resources on adding computational power to individual nodes, so an individual job can get done faster. Scaling Deeply - spending resources on understanding and optimizing an application, so each job can get done with less computational power.
Each of these can be thought of as an orthogonal axis, with its own curve of diminishing returns.
Cython is used everywhere and can be compiled on the machine it's about to be used on. It also follows best practices for talking to python via FFI.
I agree that cython is a good option if you only care about CPython.
Maybe cpyext has gotten faster since then, but I think that's the state of things still.
This strikes me as a perfect use case for Rust, too: it has great compile time memory and other safety guarantees with a speed that is likely close to (or perhaps even better than "safe") C or C++.
More generally though you're being unfair by not even seriously considering the option that there could be a legitimate reason. Also you're simply being an insulting with your common sense comment.
You are being disingenuous, for pretending this is about truth in some way and you're coming across as being an asshole by insulting people.
1) python setup.py install. Read a Cython tutorial.
2) Install the Rust toolchain and learn Rust. Hope that the language hasn't changed drastically in the past month.
Hm...
But I appreciate the "asshole" ad hominem.
Given that there's an entire article about the technical merits (and success) of this approach, and Armin is someone who is known for building excellent, widely used software, well...
/me waits for pcwalton too
Sorry, writing a blog post, while it is Armin Ronacher's forte (yes, I use the full name as we are not on first-name basis like other HN cool kids are), does not disprove that this is the riskier choice and was made because the author likes Rust.
IMHO the end result is more maintainable, readable, and accessible from FFI point of view. Regarding the performance, so Go has a GC, but I'm wondering if that would affect things dramatically at all.
Here is the Ruby side FFI code: https://github.com/jondot/scatter/blob/master/lib/scatter.rb
And here's the "native" part: https://github.com/jondot/scatter/tree/master/ext
Every now and then I keep looking at Rust and how it can integrate with higher level languages, the last time I really wanted OpenCV to work well with Rust. I think that's a big selling point. So far, to me, it's not perfect yet but it may get there.
From a pragmatic point of view, I imagine Sentry getting more bang for a buck with Go as there would be less wheels to invent from an ecosystem POV, and from a maintenance POV it would be closer to Python. But that wouldn't advance any of the Rust ecosystem at all, and we do need that as a collective.
1. There's a bunch of objects that need to pass down the ffi boundaries py->go 2. compute 3. There's a bunch of objects that need to pass up the ffi boundaries go->py 4. Python now continues as usual with a bunch of processed objects
In that case, yes this would be a problem. The way I'd resolve it is by planning the ffi boundaries accordingly. I'd make python do as less as possible, and pass just declarative "instructions" to go. In this case where's the file location and where's the sourcemap file location (and perhaps where to dump output to if that's the case). And go doing as much work as possible to make sure there's only a minimal number of objects passed back if any.
It may _feel_ like a hack but ultimately its the same approach if you were to make a "sourcemap server" making python code communicate with it over RPC.
If this is not the problem then I'd love an example of what you meant
You can look at the library in question. An object gets created in Rust but the ownership of that object is held in Python. When the Python GC runs we clean up the Rust object.
> It may _feel_ like a hack but ultimately its the same approach if you were to make a "sourcemap server" making python code communicate with it over RPC.
Sure, but that significantly complicates the problem. To the point in fact where I question if the Go solution makes any sense at all because it takes away the advantage you have where you can just drop an extension module in without much work. Once you need to restructure your system to be message based you might as well go in and run a separate process and use a unix pipe to communicate. We used to do that for a few things like our debug symbol symbolication and the downsides are just too big.
Sidenote: I wonder how improving Python performance with D fares considering it links up to C pretty nicely.
> In that case, your requirements to the language are pretty harsh: it must not have an invasive runtime, must not have a GC, and must support the C ABI. Right now, the only languages I think that fit this are C, C++, and Rust.
You can opt-out of D's GC. Turn on `-vgc` flag during compilation and replace GC'd code with non-GC'd code and you are not using any GC.
No, it doesn't make a lot of sense.
Is there a reason the Rust-exported functions aren't marked with `extern "C"`?
I do know that if you bind to a function that is linked in that you need it.
https://internals.rust-lang.org/t/precise-semantics-of-no-ma...
Regarding the declarations: this would only be necessary if processed in C++ context, to change the name mangling/linkage features of the declarations. For portability sometimes authors hide these behind "ifdef __cplusplus" barriers, but it's not really critical here.
Regarding the definitions: "Exposing a C ABI in Rust" from the article describes this in detail. For the most part, "#[no_mangle]" has the same effect that "extern "C"" has on linkage/mangling in C++.
#[no_mangle]
pub extern unsafe fn lsm_view_from_json(bytes: *const u8, len: c_uint,
err_out: *mut CError) -> *mut View
{
...
}As far as I know, `#[no_mangle]` disables name-mangling but doesn't change the ABI of a function. That's what `extern "C"` is for in Rust -- to declare a function with the C ABI. You can have a Rust ABI function with an unmangled name (what it looks is done in the post) and you can have a C ABI function in Rust with a mangled name (by using `extern "C"` but not `#[no_mangle]` -- for example for C callbacks).
Based on my limited understanding of C++, `extern "c"` in C++ is equivalent to using both `#[no_mangle]` and `pub extern "C"` in Rust. I would guess that much of the time failing to specify a C ABI would work out fine unless you try to accept or pass non-FFI types (enums, references, etc) but I'm not sure.
It's confusing as hell. There was a thread about the mixed up semantics somewhat recently on the internals forum: https://internals.rust-lang.org/t/no-no-mangle/3973 (edit: and also https://internals.rust-lang.org/t/precise-semantics-of-no-ma...).
If you look at the LLVM IR generated from https://is.gd/Hfup3X you can see that the extern "C" fn differs in that it's given a `nounwind` attribute, among other things.
Thanks for the tip btw I think this means I have a bug in my code. ;)
typedef struct { int x[500]; } bigthing;
bigthing MyFunction();
and a program using that library: int main() {
bigthing x = MyFunction();
}
C has no choice but to have MyFunction allocate a couple kilobytes on the stack and have main copy the structure from the stack to where it should eventually go. The equivalent Rust code, however, can pass the address of the object x to MyFunction, as if it were actually void MyFunction(bigstruct &output).If you have a #[no_mangle] but not extern function in Rust, the Rust compiler will generate code for that function that looks for this secret by-reference argument and fills it in. When you call it from C (or something that calls functions in a C-like manner, like Python's cffi), it won't be setting up the call like that at all, and it will expect to read the result off the stack like a C function would have done.
(I don't think there's a lot of use for #[no_mangle] without extern. The best I can think of is that, if you have a Rust application that dynamically loads a Rust library and runs a function from it, you don't have access to the mangling algorithm, once you find the function you'll call it according to the Rust ABI. But even that is risky since the Rust ABI can change between compiler versions; you're still better off shoveling things through the C ABI.)
What sort of C implementation (ABI, compiler, etc.) are you thinking of here? gcc (x86, x86-64, ARM) is perfectly capable of doing the exact optimization that you describe Rust being able to do.
I guess the trick here is that "C" really means "platform ABI" and isn't inherently about a language or a compiler.
Why not just deserialize all the source maps ahead of time and just store/retrieve them as msgpack objects?
Per this python serialization speed comparison, msgpack is ~ 10X faster than json. So you get the same speed up, but no Rust.
Why not just deserialize all the source maps ahead of time and just store/retrieve them via cPickle? Wouldn't that get you almost the same results without having to learn and support a second language (Rust, in this case)?
[Edit]
cPickle is slower than JSON, but browsing the interwebs it seems that marshal can be 2X faster than JSON and 4X faster than cPickle.
- It is not fixing python's performance.
- The performance improvement has very little to do with the choice of Rust.
They are fixing a case of Python performance being a problem in the context of their needs, and the way they solved it was with Rust (and there's no implication that it had to be Rust in the title).
It's not wrong, it's just vague.
Doesn't make Python look good.