I came back to it for a second try a month ago, and was blown away. PyO3 and maturin (especially maturin) are stunningly straightforward. I actually find maturin easier to use than “normal” Python packaging tools, to the extent such a thing even exists.
I really hope we will see a move towards rust, and away from extension modules written in C/C++/Fortran, just because the development experience is so much more approachable. Which is not something I thought I would say about Rust!
I mean, writ large, the answer is "The availability of the ecosystem they need to use daily in Rust, without having to rewrite it themselves." That is more than just a torch replacement, so this won't do it for data scientists, even if Python is ... painful in a lot of ways to productionize or even use.
Maybe it'll get some started on that ecosystem, though, so that a few years from now you see larger movement.
The uncomfortable fact is that you really don’t need thread safety or static typing to prototype and build models effectively.
Prototyping and training doesn't pay the bills. (unless you're Hugging Face, lol)
I think dynamic typing is the missing piece.
Rust's safety guarantees are great and I like the language, but the type system is a barrier for learners, especially people who aren't full-time devs. Ease of experimentation and flexibility are key for this use case.
This surprises me. Hasn't everyone working with data (even at an "enthusiast" level) had to understand data types in the context of databases?
Plus, a surprisingly enormous fraction of the NoSQL hysteria was that “you don’t even need a schema!”
The problem with this is that you have a schema regardless. In "schemaless" systems, you still have a schema. You just don't have it written down anywhere. And because its not written down, the database doesn't enforce your schema. So its very common for nosql databases to have records with missing fields or subtly corrupted data.
That said, there's a reason why nosql became popular. I think a large part of it is that SQL databases are too inflexible. You have to define the entire schema ahead of time, and schema evolution is way too difficult. When I'm writing software, I don't define all of my columns ahead of time. I want to add and change stuff as I go.
I wish databases were more flexible. Like, let me just add arbitrary extra fields to my records. Let the database recommend promoting a dynamic field to a static field. ("Hey, all your records have a name: string field. Wanna make that required?").
I love programming languages with similar features. I love typescript for prototyping because I can start with the type any (or wildcard fields in my interfaces - like [k: string]: any). And then I can formalize my types later when I understand my software better.
I don't think schemas are a mistake for beginners. I think the opposite - they're amazing guard rails.
But I agree with other commenters here. Rust's types include information about ownership, lifetime, containment (Box / Rc), traits and the stored data type all in one go. Its terrifying to beginners and exhausting for experts to use. I've been writing rust for years, and the Future trait still scares me.
It depends on the generic_const_exprs feature which is still, to quote, "highly experimental":
https://github.com/rust-lang/rust/issues/76560
Definitely not for production use, but it gives a flavor for where things can head in the medium term, and it's .. it's nice. You could imagine future type support allowing even more inference for some intermediate shapes, of course, but even what it has now is really nice. Like this cute little convnet example:
https://github.com/coreylowman/dfdx/blob/main/examples/night...
Or even serde::Value(s)? Its not that everything has to be a type
Python is not popular because it is "easy" (far from it if you have to maintain it). But because it is easy to create executable code that does what you want, without having to have a background in coding.
If you refer to lifetimes and borrow checker, those are different.
If they didn't migrate to Julia they certainly won't migrate to Rust.
With intensive Python processing being mostly C anyway, most data processing doesn't really benefit much from a Rust implementation. The actual native code implementation can probably be a lot better if that part switches over to Rust, with all of Rust's safeties and guarantees, but I don't think data science is going to switch to Rust for gluing together their actual calculations any time soon.
This library is targeting production, a much smaller subset of engineers. See the 'why candle' section: https://github.com/huggingface/candle#why-candle. So you could use PyTorch to build models and Candle to serve them.
The two things that keep Python as the incumbent in this space is A.) It gets out of the way, sometimes to a fault, such that you can kinda write it without knowing what you're doing. This is a feature for most people, not a bug.
And B.) All the math/science libraries that already exist for it.
You can solve B, but until you solve A, Python stays put I think.
I think unseating Python as the lingua franca will takes as long as it has taken to unseat MATLAB in engineering. 20+ years...
Of course other things will also exist, and all languages will have decent ML support (especially for inference).
REPLit for Rust: https://replit.com/languages/rust
For general REPL and rust appears there are a few options.
I don't think getting rid of Python completely will happen. But current approaches _typically_ looks something like:
- Jupyter Notebook for custom parts of EDA
- Jupyter Notebook for testing new libraries or learning new methods
- Pipeline (replaceable with Rust) for standard EDA
- Pipeline (replaceable with Rust) for data quality
- Application/pipeline for product development