Where Rust really shines
manishearth.github.io
manishearth.github.io
I just spent an entire week working on a chatbot in Rust, and I found the steep learning curve to be very challenging. While Rust seems to be a big win for teams that are accustomed to C++/C development, it does seem to be very demanding for web devs who always work with a garbage collector.
In particular, I was trying to use Iron to write an http server that could receive web hooks. Iron, quite sensibly, handles each incoming request to a server in a separate thread. However, as a beginner it's very challenging to figure out how to persist data between requests without writing to disk or persisting to a db. Mutating state across threads is hard. Currently, there's also no version of channels that are single producer, multiple consumer. I'm sure as the ecosystem develops, more examples will make this easier, but for now it's surprisingly difficult to figure out.
Fortunatey, the #rust IRC channel supported me all the way. It's a very friendly and active community, but without their support I would have probably given up.
I look forward to seeing more posts at this level of detail help those of us struggling to learn Rust understand how the experts use Rust.
Yeah, a lot of people are excited about Rust for webdev, but there is a steep learning curve and it's quite different from the languages most are used to. For webdev I personally think Go will make it, it feels a lot like Python/Ruby when programming, even though it's compiled and probably faster.
I'm hoping as time goes on to include more stuff in the docs that's about generic systems stuff, rather than 100% Rust specific, exactly for this reason. The stack & heap chapter of the book (which is upcoming) is part of that. But, I gotta finish off the Rust-specific stuff first...
I agree with wycats: Rust is going to bring systems programming to a lot of people who haven't done systems programming before, just like Node brought a whole new group of people to server-side code. It's our job to give you the support you need to be successful at it.
Thing is about web dev is that (1) it's a get-things-done playing field and (2) primarily just data delivery. The idea that you would want a language that doesn't have garbage collection is just silly. Why? What complex operations are you doing in each request/response cycle that demands this sort of computational horsepower? And if you are doing said complex operations, maybe you should rethink how you're delivering that data and creating it? GitHub is huge and runs on Ruby. Surely, you can't be wanting to use Rust to improve your performance are you?
> Mutating state across threads is hard
Yep, so maybe you should target a language / framework that abstracts this away from you? I wouldn't say anything usually but you say it's "surprisingly difficult to figure out." Come on really.. programming languages are not created just for web devs...
I don't know. If you're trying to be productive AND change your toolset, it sounds like you just want a typed language. I wouldn't put bet too much of your time-and-focus-chips on Rust becoming a language for web developers. Do you see many people using C++ to write their servers? Nope. And don't take my message the wrong way. It really just bewilders me because I see a comment like "Yea I want to help people like you!" and I'm like Huh??
There may be some difference in degree, though.
In more complex situations, it's generally worth thinking which lifetime should be bound together for the first few functions and structs. After that point just trust the compiler.
It's not too hard to get a mental model of lifetimes, really, and once you have it you can look at lifetime parameters and figure out what they mean and how they work with each other. You're anyway implicitly thinking of them as scopes. I usually just ignore lifetime parameters -- they're something you learn to gloss over whilst reading Rust code, and I only really read them (and try to understand them) when I have lifetime errors. Sometimes the lifetime error is due to something unfixable, in which case the compiler will often lead me through a wild goose chase saladifying the entire codebase, or the lifetime error is fixable, in which case the compiler's suggestions usually work. Sometimes they don't, but in those situations you can still fix it by trying to understand the why of it (like how I did in this post, though in the case of this post the entire "understand the why of it" was an afterthought) because things worked.
I'd love to see a kind of rust ownership tutorial in which you are asked to address a simple CS problem where these patterns occur in Rust. For example, many seem to find it hard to write a factory type. The next problem could be a doubly linked list and so on.
Which is why it works. But isn't it nice when the compiler is aware of the "use-after-free risk scale" for your data structures?
Just having a lifetime doesn't mean that it's safe to put any random borrowed data in. In C++ we could have a single pointer to a very long-lived struct ("SubstructureFields"), and wish to introduce another struct ("FieldInfo") which contains a pointer to something that is shorter lived. Note that in large codebases knowing which is "longer lived" is not easy, so from the programmer's perspective there are just two pointers. Assuming that "Okay, we don't have any segfaults now due to the first pointer, introducing the second FieldInfo pointer should be fine then" would be fallacious -- we might be accessing data during a period of time when the first, original pointer is alive, but the second is invalidated. Use after free.
Update: You don't have to wade through a lot of code at all. All the code is taking AST parameters that are the source of these Attribute vectors by const reference and returning stuff like a P<Expr> and P<Item> and the like, AST stuff that by its nature doesn't have references with tricky lifetime dependencies to other far-flung AST stuff. So in C++ you can see right from the type signatures (and some basic institutional knowledge) that you're OK.
You can see that the lifetimes are not long or dangerous. SubstructureFields is used in Substructure. A simple grep for that type shows a bunch of functions that take a Substructure by const reference and return a P<Expr>. There is nothing holding wiggly little references to Substructure objects or the like.
You don't need to debate with me the merits of replacing visual analysis with sound analysis. (If I have a negative opinion of Rust on the matter, it's that it's not good enough at that.) My beef here is with the way the blog article overstates the case, saying you just wouldn't do this sort of thing or this specific thing. You so would make temporary references deep into an AST while expanding derived implementations.
Yes, in this situation owned Expr pointers are returned, and fortunately I know that Expr contained no borrowed data (with a lack of lifetime annotation that is easy, but in C++ you might have to check)
But there could be other pointers being passed around internally (eg replacing the old pointer from a field of an argument with a new one, one which actually is short lived or is being iterated over. Just because a pointer isn't returned doesn't mean we're safe. With const we get some safety until interior mutability and casts come into the picture -- as I said before, consting a complex codebase is not easy unless you keep switching back and forth.
SubstructureFields could have had a lifetime parameter of something that was supposed to live longer, we can't be easily sure.
This trick doesn't let me track the flow of a program to find out which portions of code will be run while my pointer is active. Tracking the flow only happens in Rust.
Const correctness has nothing to do with it. It does not enforce reference safety in C++, and the Rust compiler does not use "const correctness" to enforce proper memory management.
Unless you're trying to say that it's not a formal proof of correctness, but that's not an argument that engages with the claim the blog article made about whether this is something you would or wouldn't, or couldn't do, in C++.
I think the nature of the 1.0 release and the train model is not very easy to communicate: the Rust team hopes to commit strongly to semver, so "stable Rust" is only that which they are comfortable committing to providing backwards compatibility for indefinitely. It's really pretty fine to be on the nightly track if you don't mind occasionally having to change a method call or something when an unstable API you use change.
Not trying to dismiss what you've said - I just think that the strength with which you expressed it ("Rust is still barely usable") was not accurate & that if you ask in the IRC channel or users.rust-lang.org, there is likely a no-or-little-hassle workaround for whatever feature you wanted.
>In a language like C++ there’s only once choice in this situation; that is to clone the vector.
What about a shared_ptr to the vector?
Initially this gets annoying, but as time passes one realizes how awesome this is. It prevents iterator invalidation (and vector-move-invalidation), for one (it's also necessary to get memory safety for our ADT enums since you can otherwise change the variant and invalidate a reference to its contents). In general you realize that in a large codebase, it's almost as bad as a multithreaded situation -- you don't know what objects are being modified by the methods being called and the code is too large to easily figure it out.
There's certainly more than one choice, though.
These things protect against one part of the problem - deletion of the vector, but not the other part - mutation of the vector, causing a new internal allocation, and leaving references to any elements of the previous vector invalid.
It's possible to make sure this isn't happening in C++ if a) your code base is small enough and b) you are careful, but the point of the article is that with Rust, you can make the changes and be confident that the compiler will fail if you do anything dangerous.
With C++, you can make the changes and through careful checking be reasonably confident everything works (perhaps only to find out later that you were wrong). It's very different from having the compiler verify that you are not doing something unsafe.
In Rust, when you have an immutable reference to a memory location, you know for certain that it won't be modified by distant code as long as you hold the reference. This guarantee also allows more compiler optimizations, because v[5] has the same value after a function call as before.
Also, iterator invalidation.
Also, runtime overhead of shared_ptr.
(I added a footnote to the post to explain the issues with shared pointers)
struct FieldInfo {
//
const std::vector<ast::Attribute>& attrs;
}It is true it will not check at compilation that you access a const reference of a destroyed object, but to be honest with RAII this is not really a problem as the lifetime of your objects should be pretty clear, by construction.
In debug mode your program will quickly assert if you access a destroyed STL container.
To be honest I think in this case:
- either you want to snapshot the value and therefore you should clone the container (almost no performance cost for small containers)
- you want to access the current value but then it means you know about the lifetime.
Given the potentially-unbounded consequences of triggering UB in C++, "pretty clear" seems like not enough.
I find that a bit worrying - while it might well be an example of what's technically possible in some cases of using Rust to prevent segfault-type crashes, it doesn't mean there aren't now logic errors in the code which would do something just as bad. I guess it depends on how well you know the code and the context in which it's being used.
Yes, it does. He wants to have immutable access to some data, the Rust borrow checker's rules provide certainty that the data is not deleted or mutated while his code is accessing it.
What could happen that is "just as bad"? Data being referenced cannot be changed.
I was more talking about the overall not caring theme. I'd hope that's only because they know the code so can make those sort of assumptions.
Are you saying Rust can find and prevent at compile time logic errors as well?
Rust allows you to not care about more things than C/C++ despite compiling to a comparable binary. In particular, Rust makes refactoring and extending a code base much more care-free than many other languages (including high level languages) because of the strong type and memory safety guarantees of its semantics.
EDIT (re: your edit) - No one said anything about preventing arbitrary logic errors, just that what Manishearth did will not introduce logic errors because of Rust's guarantees. This could not be the data that Manishearth wants, but the compiler guarantees that Manishearth cannot change the data or delete the data and that the data is not changed or deleted while he is referencing it. That is what is so great.
That was my original point regarding logic errors and not caring - yes, it may well be that Rust gives you much more confidence that you don't do invalid/wrong things. But that excludes logic errors (it may well reduce the possibility of logic errors, but it doesn't prevent them), and there's still room in large code bases for things to break even with very good static type checking.
> I was more talking about the overall not caring theme.
Not having to care about memory management gives me more mental energy for caring about the other parts of my program. The lack of anxiety about memory semantics is one of the most compelling reasons for me to use Rust. > Are you saying Rust can find and prevent at compile
> time logic errors as well?
In some ways it certainly can. Its type system guarantees at compile time that your code cannot contain data races, making massive concurrency and parallelism much easier to introduce. You can also push the type system to do things like check units for you, such as is done in Servo: https://blog.mozilla.org/research/2014/06/23/static-checking...And man, functional languages. It's pretty much 100% logic bugs there. I worry most about those sort of programs.
Different programming platforms have different "incidental complexity" causing bugs not related to actual program logic, but merely how it is put together, so to speak. In particular, Rust seems objectively better than C++ in this regard.
I was just holding on to a reference for an API to consume.
Of course, in more complicated situations you need to worry about logic errors. But in this case I was pretty sure I didn't have anything to worry about.
Still you never have to worry about memory unsafety, which is a huge cognitive load off my mind. Similar to how you don't have to worry about some types of memory unsafety in languages like Java. Or how most programmers don't worry about cache coherence since the processor handles it. Rust just raises the bar. If you know what it can and can't prevent, you can be carefree about many things.
It's like saying that driving your car off a cliff is always a bad idea even when you've acquired a car that can fly.
Then again, contracts about who can modify what data are also part of any public API; historically, there have been plenty of (quite decent) APIs that have returned pointers to internal structures that have limited lifetimes, in languages like C that don't enforce lifetimes at all.
For all use-cases that are possible this kind of API it would be safe to use a const reference to a vector it C++ too. Of course with the difference that in Rust the correct use can be statically verified.
But I think for a really flexible API that allows parallel modification you should go for a copy or some kind of immutable-vector oder copy-on-writer-vector type in both languages.
(Besides, all these were internal functions that wouldn't be used as an API)
You don't always want parallel modification.
But it limits how the API (especially on provider site) can be used in the future. If there comes up a need that that the data might need to be mutated at one point in future you might need to change it to an other interface in both of the languages and then change probably lot's of consumers according to it.
If it's only an internal API that might be ok - for a public API it would be to inflexible for my taste.
Generally to add mutation you might add another method in the situation you describe, but I can't think of such a situation coming up in a well-designed Rust API. (You're treating the API as if it would be designed the C++ way -- it wouldn't)
FWIW most libraries do use a COW type in many places to get this flexibility. But in most situations you can just return mutable or immutable slices.
It's about that this API design causes close coupling between the API and the internal representation of the API. If you change your internal software structure and the data is stored in another way or you choose the make your library concurrent internally where different parts could mutate the vector while the user might request the same data in parallel then returning the slice would no longer be possible. And so you would have to change your external API and all consumers of the API only because you wanted to change internal things.
But as I said: Depends on the use-case.
In that case you wouldn't want to return a slice, because it would be dangerous. The compiler would prevent you from invalidating an API use case. I don't see a problem here.
This platform advocacy comes from what business needs (ability to extort money by gatekeeping), not what's technically sensible nor useful for the people.
Plus, having a GC for this would be quite heavy. The codebase was a compiler -- performance is necessary since a compiler's going to be slow anyway.