Announcing Rust 1.22
blog.rust-lang.org
blog.rust-lang.org
Also - any plans for Streaming?
I'm a big Scala user and fan (Haskell of course too) and would like to try out Rust but I really don't see myself going back to structural/object oriented-programming
[1]: http://smallcultfollowing.com/babysteps/blog/2016/11/02/asso...
[2]: https://github.com/rust-lang/rfcs/pull/1598
[3]: http://smallcultfollowing.com/babysteps/blog/2017/01/26/lowe...
It is unlikely you will ever see non-zero cost abstractions in idiomatic Rust.
That said, at some point, we'll be getting ATC, which is equivalent to HKT in many cases, but feels more Rusty.
- HKT: Higher-Kinded Types
- ATC: Alternative Type Constructors
From what I remember from the theory, higher-kinded types lead to a genuinely more difficult type inference problem (going from easy first-order unification to undecidable higher-order unification), while associated types are simply existential types (and in particular no more difficult than higher-rank polymorphism).
The feature under discussion is "associated type constructors". Rust already has associated types in traits (I didn't know that part and was confused), and what this feature adds is that it allows us to associate a first-order type constructor to a trait.
Since the type constructor is first-order, and first-order type constructors are already present in the base language in the form of generic types, the implementation is simplified to the point that it can reuse the existing infrarstructure for type inference.
---
Apart from that, the reason for this implementation choice seems to be that it's required for precise lifetime management. Almost all datastructures in rust seem to be parameterized with a lifetime argument, even if they have no further type parameters. Since there is no such thing as a second-order lifetime (i.e., a "lifetime constructor" T : (lifetime -> lifetime) -> lifetime), first-order type constructors are enough to handle all issues that pop up because of lifetime management.
---
That actually seems like a very pragmatic design. The only problem I had while reading this RFC is that the combination of "multiple-inheritance" in traits with their built in namespacing leads to some really ugly syntax, e.g., "<T as Foo>::Bar<'a, 'static>;". Is this already idiomatic rust?
There's several more things that confuse me in this RFC, but this is probably the wrong place to discuss these things. Incidentally, what is the right place to talk about this?
It's only used for disambiguation cases, extremely rarely. I've been doing Rust almost five years and I think I may have written <T as Foo> once.
> what is the right place to talk about this?
https://internals.rust-lang.org/ is the best place. Probably going to be pretty slow-going until next week, due to the holiday today.
Re "extra", I'm not sure what you mean - does it mean that the abstraction is implemented efficiently, but still has the cost that is usually associated with that abstraction? For example, dynamic dispatch is as fast as a DIY dynamic dispatch in assembly?
The borrow checker, AFAIU, has a cost in that it forbids some (more efficient) shared data accesses that it can't prove safe but still are.
Yes, zero cost means you could not implement it in a more efficient way by hand.
> The borrow checker, AFAIU, has a cost in that it forbids some (more efficient) shared data accesses that it can't prove safe but still are.
This is not what it meant with "cost" in this context. It is zero cost because it only runs at compile time. That it rejects some valid programs is irrelevant, because it is zero cost for the programs it accepts.
Safe Rust has runtime costs as well (mostly bound checks).
For an illustration of a safety-related runtime cost that you almost always have to pay, Uft8 validation of strings is a better example IMHO.
This means it's in 1.23, which'll be released in 6 weeks.
You should note that it's not as simple as that, though. Depending on the context you want to run WASM you might need to do additional steps. E.g., when running this in a browser, you might not be able to use the std lib to spawn threads just yet.
let mut x = 2;
let y = &8;
// this didn't work, but now does
x += y;
Why is it desirable that that work? It seems like it will mask mistakes and make types more confusing.People who say Rust is hard to learn have always confused me. I learned programming in this order: scheme, C, C++, Haskell, Rust. I believe that if you go in another order it may make things more difficult, but some of these difficulties are inherent. To me, a reference to a thing is inherently different from the thing, and a systems language should make that clear. I didn't even like auto-deref, I felt like there should be an operator similar to
thing->method()
in C++ that can be chained so that references and indirection would always be syntactically clear and distinct, but the ergonomics on that won out.(proceeds to list a quite rare and outlier-ish learning sequence, and all in non-scripting languages to boot).
Well, of course those people confused you. What you, who learned Scheme, C++ and Haskell before going to Rust, have in common with someone who only knows C or Java or only got his start in some scripting languages, so as to be able to relate with their "Rust is hard" experience?
Folks always say this but never really clarify why this should be the case.
When doing low lever unsafe stuff Rust does indeed make it very clear if something is a pointer or not, and avoids weird coercions (even C doesn't do well here!)
But for normal code, does it really matter if something is a pointer or not? In C++ already with by-ref and move values and stuff you have implicit pointers in many places that don't behave as pointers.
Rust never goes the way of adding implicit allocations or implicit indirection; it only goes in the way of stripping it, which is rarely if ever a problem except when dealing with unsafe code (where Rust is more explicit on pointers anyway).
As for confusion, both are accepted. I imagine clippy will eventually have advice on this situation. For example, clippy also warns you about taking a reference that the compiler immediatly dereferences.
I’m not sure what the long term plans are for clippy, but personally hope it’ll eventually be warnings in the official compiler.
I just wanted to point out that this is largely a misconception. Rust, or the Rust creators, do not consider pointer arithmetic to be fundamentally `unsafe`, and the only reason `.offset()` is `unsafe` to begin with is due to optimization concerns[0]. `.wrapping_offset()` exists and is marked safe, despite achieving the exactly same thing as `.offset()` for the majority of scenarios[1], and more-over you can cast any pointer to an integer using only safe code, at which point you can perform all of your arithmetic on the integer and then cast it back also using only safe code[2]. The bottom line is that things like pointer arithmetic are generally not considered unsafe because the operation has defined semantics for what happens. It is only the point where you attempt to use the value, the dereference, that is `unsafe`, as that may blow-up if the value is an invalid memory location.
[0] https://github.com/rust-lang/rust/issues/33813
[1] https://play.rust-lang.org/?gist=c9c11545fac3a45550e6810cb57... An example that uses `.wrapping_offset` without `unsafe` to produce a segfault.
[2] https://doc.rust-lang.org/book/first-edition/casting-between...
My point was more that there is a difference between what people consider to be "bug-likely" operations, and what Rust considers an `unsafe` operation, and because of this I think people are too quick to think the `unsafe` system will always save them and make it easy to verify their code. If you do tons of pointer arithmetic and then do one dereference at the end, you're going to only have one line of `unsafe` code, which is "good". But if that one line blows up, then the actual bug is somewhere in your safe code, not the one line of `unsafe` code, so judging safety based only on how much `unsafe` code is really not a great way to do things.
And I don't say this as a theory about what people are thinking. This paper[0] got passed around a bit a while back (Most only Reddit, I can't seem to find a HN page that got very popular), and they quite literally create an `Address` object that exposes a safe `plus` function (for pointer arithmetic) and an `unsafe` `load` function for dereferencing. And then they just do a pretty much straight conversion of the C code, replacing all the pointers with `Address` objects, replacing all of the pointer arithmetic with `plus`s, and replacing all of the dereferences with `load`s. And then they claim it is tons safer then the C code because it uses so little `unsafe` code, despite the fact that it can easily have the exact same bugs the C code can have if there are any bugs in their pointer arithmetic. So, in this situation, the amount of `unsafe` code really doesn't actually matter because it doesn't really mean anything about the number of bugs in the program. All it indicates is that "these are the spots where the program could blow-up", which you could probably already figure out fairly easy from looking at C code.
And don't get the wrong idea, I do still like Rust and I think it has some great features in it (And some misfeatures, but that's true of every language). But, at the very least, I think the `unsafe` system leaves a lot to be desired and things like the tutorials give the wrong impression about what `unsafe` actually indicates.
[0] https://pdfs.semanticscholar.org/4dc9/61a6922eae5346c593eb18...
There's obviously not a way to enforce that in the compiler, but that's because unsafe is inherently a way to extend what the compiler can enforce. It's like you're adding code to the compiler.
Assuming that's done correctly (as e.g. the stdlib is in the process of being proven to do) then any client safe code is truly known not to have any memory safety problems.
Also, define 'correct'. What does it mean for a `unsafe` block to be 'correct'?
My point is largely that for an `unsafe` block to be correct, it relies on other safe code to also be correct in ways the compiler can't verify. So just because your `unsafe` blocks are small and infrequent doesn't say anything about the quality or "buggyness" of your code as a whole if those `unsafe` blocks are dependent on largely the entire rest of the safe code.
This is true; what I'm getting at is that you shouldn't allow your unsafe code to depend on "largely the entire rest of the safe code."
Unsafe code should be used together with the visibility system, the borrow checker, and the trait system, so that it only depends on a very small amount of code within the same file or even the same function.
Then the vast majority of the program can be written in safe code that cannot violate the invariants of the unsafe code.
With all that said, while this is a separate topic, I think that the complete lack of a spec for lots of details surrounding `unsafe` makes it still basically impossible to do correctly. I think Rust really needs some type of language spec, but it doesn't look like any work is being done on that front [0].
[0] https://github.com/rust-lang/rust/issues/30500 - I apologize for this link being a bit of a rabbit-hole, but I think the whole thing is a very interesting read. It continues on here (https://github.com/rust-lang/rfcs/issues/1447), and then to a repo for a Rust Memory Model specification here (https://github.com/nikomatsakis/rust-memory-model). Unfortunately, as you can probably tell, those conversions started in Dec 2015, and any real work on this front seems to have stalled almost a year ago.
let mut x = 2;
let y = &8;
x = x + y;
What mistakes could this be masking? I can't think of any.Mistakes in terms of understanding the code and writing what you think you're writing. I've only just begun learning Rust, and one of the most difficult things has been figuring out what types my variables have when everything is implicit and based on type inference. Variables that act both as pointers and non-pointers seem pretty confusing. It's like temporarily and invisibly turning off parts of the type-checker.
ETA: To me it feels a bit like the icky type coercion magic some languages have that lets you write code that kind of works even though you don't understand what you're doing. I don't know Rust though, so this may be totally different.
Here is one C++ example, without looking for the signature of f(), there's no way to tell the answer. So essentially two pieces of identical code with identical input values and no external side effects can still give you different output. Ah C++, the garbage that you are.
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
So why is Rust any better in this regard? Because at the end of the day, Rust deals with actual concrete types and values, the only thing references are good for is being references: they sometimes save you some copying and let you refer to stuff and that's pretty much it. So the "ref1 == ref2" operation is a deep equality check and not a "does ref1 point to the same object as ref2" check. Because sometimes objects in different places in memory can be semantically equal. So it's okay for obj1 to mean obj1, &obj1 to mean obj1, &&obj1 to mean obj1, etc...If you need raw pointers like you need in C, because you know the objects you want to compare are singletons, then you can always cast using
&obj as *const Obj
or &mut obj as *mut Obj
But that's an escape hatch.BTW, this auto-dereferencing you're experiencing is part of the Deref trait if you ever want to overload it. But that's, in my own opinion, an escape hatch as well.
All in all, unless you're doing very specific things, try to write your code in high-level semantics (meaning using these deep-equality rather than pointer-equality semantics), and then benchmark and find out which parts are hurting your performance. Rust allows you to do that and still get between very-reasonable-and-very-good speed.
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
if you really want to make sure that x is not changed by f, declare it const. Sure, f can cast the constness away, but then again it could also walk up the stack in C and trash your stack frame anyway. It is UB in either case.The typesystem in C++ can help protect against Murphy, but (unfortunately) not Machiavelli.
int x[] = {5};
f(x);
// WHAT IS THE VALUE OF X HERE?
see, it works in C as well.(1) They are two different types and (2) it's clear to you at the calling site!
More than 200 documented use cases of undefined behavior...
Are we talking about the same language here? The language where arrays implicitly decay to pointers, where integer types get implicitly promoted all over the place, where aliasing rules implicitly define which pointers can and cannot alias, where partial initialization of a struct or array implicitly sets the other members to 0? Where the language will let you call an undeclared function and make up a prototype on the fly for you?
I also don't understand your C++ example, without side-effect why would both invocations of f() within the same scope produce a different result? I thought you wanted to criticise function overloading but you call it with an int both times so I don't see what's you're getting at.
Or maybe you meant that the two calls are in a different scope and could resolve to a different function? But you can do that in C as well in two different translation units and making two static f() implementations.
-Wall -Wextra and those issues are made clear to you. But granted, that's part of a good compiler and not part of the language.
> I also don't understand your C++ example, without side-effect why would both invocations of f() within the same scope produce a different result? I thought you wanted to criticise function overloading but you call it with an int both times so I don't see what's you're getting at.
I'm not criticizing function overloading.
> Or maybe you meant that the two calls are in a different scope and could resolve to a different function? But you can do that in C as well in two different translation units and making two static f() implementations.
Yes the two functions are different, but not necessarily in scope, they don't necessarily need to have the same names. Yet at the calling site they look exactly the same: one will modify your data without you being aware of it and the other will not, and there is no syntactic hint to differentiate them.
It is quite different.
In Rust already it is possible to proxy "pointer to a thing" as "the thing itself" in a lot of different places, making it a lot more pleasant to work with. This is a continuation of that trend.
It's pretty easy to know if most things are a pointer or not, but more importantly if it's not immediately obvious it rarely matters (i.e. you're not the one creating or consuming it; you're just sharing borrows to it). At which point sharing pointer-to-value vs pointer-to-pointer isn't very different, and Rust implicitly deals with that.
Swift has taken the observation further and is investigating ways to avoid ever having a immutable+shared-reference-to-primitive vs primitive distinction in the language, while still introducing this distinction for types where it is interesting. (e.g. reference counted classes or atomics)
But it's not a ternary operator, it's unary.
let val = foo.bar()?; will call `bar()`, return if it is None, and if it is `Some(foo)`, set `val` to `foo`.
fn try_option_some() -> Option<u8> {
if let val = Some(1) {
Some(val)
} else {
None
}
}
It's an early exit if the expression before is None/Error fn try_option_some() -> Option<u8> {
Some(1)
} fn f() -> Result<A, E>
fn g() -> Result<B, E>
Then you can write fn h() -> Result<(A,B), E> {
let x = f()?;
let y = g()?;
Ok((x, y))
}
rather than having to do something like fn h() -> Result<(A,B), E> {
match f() {
Ok(x) => {
match g() {
Ok(y) => Ok((x,y)),
Err(e) => Err(e)
}
},
Err(e) => Err(e),
}
}
(written from memory, please excuse syntax errors)What's new is that this also works with Option<A> now.
Another sane alternative here is using and_then, map, etc.
Writing from memory too:
fn h() -> Result<(A, B), E> {
f().and_then(|x| g().map(|y| (x, y)))
}
..which seems to be a different style of control flow to choose from. On one hand, I suppose one may avoid using and_then with code dealing with side effects. On another hand, ? seems restricted to being used inside functions that have to return a a single type (like Result) they operate on.> functions that have to return a single type
Well, all functions have to return a single type, though that type may be a composite type, like a tuple.
Less pedantically, `?` lets you (well, will let you) unwrap values and convert their error cases between each other. So once the next round of stabilization happens, if you have a function that returns Results, you can mix ? on Options and Results in the body, and vice versa. And it can be extended for other types too. Basically, it's an early return "protocol" if you will.
do_this(x).ok_or(CustomError::ThisError)?;If you added extra code to manipulate things then that single if would still be enough, just place everything that happens after the extraction inside the success branch of the if statement.
You're missing a Some(val) at the end. Now you could remove the `if let val` as well to make it compile, but then that obfuscates what the ? operator was doing there.
I think it would have been much clearer to have 1 function which took a parameter of type Option<u8>, then did something with it after extracting with ?, then call it twice, once with Some(1), once with None. In addition I'd probably do something with val (e.g. add 1). And to prevent easy conversion to the if statement you could make the early return explicit by doing the extraction somewhere in a loop (so that you can't get at it with nested if statements).
I mean, this isn't hard to check, try compiling it.
I was saying it's misleading not only because it's wrong, but also because even if something is equivalent if you're trying to show how a feature works you want equivalent code that cleanly maps to the original; and isn't further reduced.
fn try_option_some(param: Option<u8>) -> Option<u8> {
if let Some(val) = param {
Some(val + 1)
} else {
None
}
}
fn main() {
print!("{:?}\n", try_option_some(None));
print!("{:?}\n", try_option_some(Some(2)));
}> if-else conditionals are expressions, and, all branches must return the same type.
It's used for shortcut error handling: when for instance you receive an error result from opening a file, you usually want to also return with an error. It was used so often that the ? operator was introduced as a shortcut. Now the operator is being extended to Option<T>, which works similar to Result<T, ()>.
And that includes a link to the 1.13 announcement with a detailed explanation. https://blog.rust-lang.org/2016/11/10/Rust-1.13.html
> Was more boilerplate necessary in the example before?
Somewhat yes, the original use case is Result types (and the try! macro):
let v = func()?;
expands to let v = match func() {
Ok(v) => v,
Err(e) => return Err(From::from(e))
};
that is it "unwraps" a successful result, and barring one it directly returns from the current function (having possibly converted the error). Before 1.22, doing that with an Option would be: let v = match func() {
Some(v) => v,
None => return None
};
or let v = if let Some(v) { v } else { return None };
after 1.22, this becomes let v = func()?;
as well.edit: note that this is quite different from the behaviour of the ? operator in languages like C# or Swift where it compounds into "null-safe" operators, the most famous being the null-safe navigation operator "?."[0], these are closer to the monadic bind, which in Rust is called "and_then".
fn try_option_some() -> Option<u8> {
let val = Some(1)?;
Some(val)
}
fn try_option_none() -> Option<u8> {
let val = None?;
Some(val)
}
fn main() {
assert_eq!(try_option_some(), Some(1));
assert_eq!(try_option_none(), None);
let v: Option<u8> = None;
assert_eq!(v, Some(None));
}
One other thing troubling me here - shouldn't try_option_none() return `None` and not `Some(None)`?Looking at the diff for the commit, there's now a NoneError? Is there any documentation updates to look at how this works? Because AFAIK, now all matches will have to explicitly be updated to check for NoneError, otherwise they won't compile (since matches are not complete)?
= note: expected type `std::option::Option<u8>`
found type `std::option::Option<std::option::Option<_>>`
The return types of those functions are Option<u8>, but Some(None) is an Option<Option<_>>.> there's now a NoneError?
It's not stable yet.
> now all matches will have to explicitly be updated to check for NoneError
NoneError only comes into play when you have a function that returns Result, but uses ? on an Option in the body. That code never compiled (and still doesn't on stable until Try is stabilized, and only if you have Result<T, NoneError> as your return type or some Err type that implements From<NoneError>, which you also can't write until it's stable.
Does that help?
And that includes a link to the 1.13 announcement with a detailed explanation.
Can't imagine why compile times haven't been good in the past with this level of detailed analysis.
Also, incremental compilation is making very good progress. More fuzzy numbers: I'm currently working on a mid-size Rust project, where the app crate has around 10k lines. `env CARGO_INCREMENTAL=1 cargo run --release` takes 2-3 seconds depending on what I changed.
But beyond that, it's ultimately about what you promise. Say something sped up compilation times 2x for most people, but slowed them 10x in a corner case. If you promise the 2x, the people for who it slowed down will be (rightfully) pretty mad.
None of these changes are as drastic as that, but still.