let mut x = 2;
let y = &8;
// this didn't work, but now does
x += y;
Why is it desirable that that work? It seems like it will mask mistakes and make types more confusing. let mut x = 2;
let y = &8;
// this didn't work, but now does
x += y;
Why is it desirable that that work? It seems like it will mask mistakes and make types more confusing. let mut x = 2;
let y = &8;
x = x + y;
What mistakes could this be masking? I can't think of any.Mistakes in terms of understanding the code and writing what you think you're writing. I've only just begun learning Rust, and one of the most difficult things has been figuring out what types my variables have when everything is implicit and based on type inference. Variables that act both as pointers and non-pointers seem pretty confusing. It's like temporarily and invisibly turning off parts of the type-checker.
ETA: To me it feels a bit like the icky type coercion magic some languages have that lets you write code that kind of works even though you don't understand what you're doing. I don't know Rust though, so this may be totally different.
Here is one C++ example, without looking for the signature of f(), there's no way to tell the answer. So essentially two pieces of identical code with identical input values and no external side effects can still give you different output. Ah C++, the garbage that you are.
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
So why is Rust any better in this regard? Because at the end of the day, Rust deals with actual concrete types and values, the only thing references are good for is being references: they sometimes save you some copying and let you refer to stuff and that's pretty much it. So the "ref1 == ref2" operation is a deep equality check and not a "does ref1 point to the same object as ref2" check. Because sometimes objects in different places in memory can be semantically equal. So it's okay for obj1 to mean obj1, &obj1 to mean obj1, &&obj1 to mean obj1, etc...If you need raw pointers like you need in C, because you know the objects you want to compare are singletons, then you can always cast using
&obj as *const Obj
or &mut obj as *mut Obj
But that's an escape hatch.BTW, this auto-dereferencing you're experiencing is part of the Deref trait if you ever want to overload it. But that's, in my own opinion, an escape hatch as well.
All in all, unless you're doing very specific things, try to write your code in high-level semantics (meaning using these deep-equality rather than pointer-equality semantics), and then benchmark and find out which parts are hurting your performance. Rust allows you to do that and still get between very-reasonable-and-very-good speed.
int x = 5;
f(x);
// WHAT IS THE VALUE OF X HERE?
if you really want to make sure that x is not changed by f, declare it const. Sure, f can cast the constness away, but then again it could also walk up the stack in C and trash your stack frame anyway. It is UB in either case.The typesystem in C++ can help protect against Murphy, but (unfortunately) not Machiavelli.
int x[] = {5};
f(x);
// WHAT IS THE VALUE OF X HERE?
see, it works in C as well.(1) They are two different types and (2) it's clear to you at the calling site!
More than 200 documented use cases of undefined behavior...
Are we talking about the same language here? The language where arrays implicitly decay to pointers, where integer types get implicitly promoted all over the place, where aliasing rules implicitly define which pointers can and cannot alias, where partial initialization of a struct or array implicitly sets the other members to 0? Where the language will let you call an undeclared function and make up a prototype on the fly for you?
I also don't understand your C++ example, without side-effect why would both invocations of f() within the same scope produce a different result? I thought you wanted to criticise function overloading but you call it with an int both times so I don't see what's you're getting at.
Or maybe you meant that the two calls are in a different scope and could resolve to a different function? But you can do that in C as well in two different translation units and making two static f() implementations.
-Wall -Wextra and those issues are made clear to you. But granted, that's part of a good compiler and not part of the language.
> I also don't understand your C++ example, without side-effect why would both invocations of f() within the same scope produce a different result? I thought you wanted to criticise function overloading but you call it with an int both times so I don't see what's you're getting at.
I'm not criticizing function overloading.
> Or maybe you meant that the two calls are in a different scope and could resolve to a different function? But you can do that in C as well in two different translation units and making two static f() implementations.
Yes the two functions are different, but not necessarily in scope, they don't necessarily need to have the same names. Yet at the calling site they look exactly the same: one will modify your data without you being aware of it and the other will not, and there is no syntactic hint to differentiate them.
It is quite different.
In Rust already it is possible to proxy "pointer to a thing" as "the thing itself" in a lot of different places, making it a lot more pleasant to work with. This is a continuation of that trend.
It's pretty easy to know if most things are a pointer or not, but more importantly if it's not immediately obvious it rarely matters (i.e. you're not the one creating or consuming it; you're just sharing borrows to it). At which point sharing pointer-to-value vs pointer-to-pointer isn't very different, and Rust implicitly deals with that.
Swift has taken the observation further and is investigating ways to avoid ever having a immutable+shared-reference-to-primitive vs primitive distinction in the language, while still introducing this distinction for types where it is interesting. (e.g. reference counted classes or atomics)
As for confusion, both are accepted. I imagine clippy will eventually have advice on this situation. For example, clippy also warns you about taking a reference that the compiler immediatly dereferences.
I’m not sure what the long term plans are for clippy, but personally hope it’ll eventually be warnings in the official compiler.
I just wanted to point out that this is largely a misconception. Rust, or the Rust creators, do not consider pointer arithmetic to be fundamentally `unsafe`, and the only reason `.offset()` is `unsafe` to begin with is due to optimization concerns[0]. `.wrapping_offset()` exists and is marked safe, despite achieving the exactly same thing as `.offset()` for the majority of scenarios[1], and more-over you can cast any pointer to an integer using only safe code, at which point you can perform all of your arithmetic on the integer and then cast it back also using only safe code[2]. The bottom line is that things like pointer arithmetic are generally not considered unsafe because the operation has defined semantics for what happens. It is only the point where you attempt to use the value, the dereference, that is `unsafe`, as that may blow-up if the value is an invalid memory location.
[0] https://github.com/rust-lang/rust/issues/33813
[1] https://play.rust-lang.org/?gist=c9c11545fac3a45550e6810cb57... An example that uses `.wrapping_offset` without `unsafe` to produce a segfault.
[2] https://doc.rust-lang.org/book/first-edition/casting-between...
My point was more that there is a difference between what people consider to be "bug-likely" operations, and what Rust considers an `unsafe` operation, and because of this I think people are too quick to think the `unsafe` system will always save them and make it easy to verify their code. If you do tons of pointer arithmetic and then do one dereference at the end, you're going to only have one line of `unsafe` code, which is "good". But if that one line blows up, then the actual bug is somewhere in your safe code, not the one line of `unsafe` code, so judging safety based only on how much `unsafe` code is really not a great way to do things.
And I don't say this as a theory about what people are thinking. This paper[0] got passed around a bit a while back (Most only Reddit, I can't seem to find a HN page that got very popular), and they quite literally create an `Address` object that exposes a safe `plus` function (for pointer arithmetic) and an `unsafe` `load` function for dereferencing. And then they just do a pretty much straight conversion of the C code, replacing all the pointers with `Address` objects, replacing all of the pointer arithmetic with `plus`s, and replacing all of the dereferences with `load`s. And then they claim it is tons safer then the C code because it uses so little `unsafe` code, despite the fact that it can easily have the exact same bugs the C code can have if there are any bugs in their pointer arithmetic. So, in this situation, the amount of `unsafe` code really doesn't actually matter because it doesn't really mean anything about the number of bugs in the program. All it indicates is that "these are the spots where the program could blow-up", which you could probably already figure out fairly easy from looking at C code.
And don't get the wrong idea, I do still like Rust and I think it has some great features in it (And some misfeatures, but that's true of every language). But, at the very least, I think the `unsafe` system leaves a lot to be desired and things like the tutorials give the wrong impression about what `unsafe` actually indicates.
[0] https://pdfs.semanticscholar.org/4dc9/61a6922eae5346c593eb18...
There's obviously not a way to enforce that in the compiler, but that's because unsafe is inherently a way to extend what the compiler can enforce. It's like you're adding code to the compiler.
Assuming that's done correctly (as e.g. the stdlib is in the process of being proven to do) then any client safe code is truly known not to have any memory safety problems.
Also, define 'correct'. What does it mean for a `unsafe` block to be 'correct'?
My point is largely that for an `unsafe` block to be correct, it relies on other safe code to also be correct in ways the compiler can't verify. So just because your `unsafe` blocks are small and infrequent doesn't say anything about the quality or "buggyness" of your code as a whole if those `unsafe` blocks are dependent on largely the entire rest of the safe code.
This is true; what I'm getting at is that you shouldn't allow your unsafe code to depend on "largely the entire rest of the safe code."
Unsafe code should be used together with the visibility system, the borrow checker, and the trait system, so that it only depends on a very small amount of code within the same file or even the same function.
Then the vast majority of the program can be written in safe code that cannot violate the invariants of the unsafe code.
With all that said, while this is a separate topic, I think that the complete lack of a spec for lots of details surrounding `unsafe` makes it still basically impossible to do correctly. I think Rust really needs some type of language spec, but it doesn't look like any work is being done on that front [0].
[0] https://github.com/rust-lang/rust/issues/30500 - I apologize for this link being a bit of a rabbit-hole, but I think the whole thing is a very interesting read. It continues on here (https://github.com/rust-lang/rfcs/issues/1447), and then to a repo for a Rust Memory Model specification here (https://github.com/nikomatsakis/rust-memory-model). Unfortunately, as you can probably tell, those conversions started in Dec 2015, and any real work on this front seems to have stalled almost a year ago.
People who say Rust is hard to learn have always confused me. I learned programming in this order: scheme, C, C++, Haskell, Rust. I believe that if you go in another order it may make things more difficult, but some of these difficulties are inherent. To me, a reference to a thing is inherently different from the thing, and a systems language should make that clear. I didn't even like auto-deref, I felt like there should be an operator similar to
thing->method()
in C++ that can be chained so that references and indirection would always be syntactically clear and distinct, but the ergonomics on that won out.(proceeds to list a quite rare and outlier-ish learning sequence, and all in non-scripting languages to boot).
Well, of course those people confused you. What you, who learned Scheme, C++ and Haskell before going to Rust, have in common with someone who only knows C or Java or only got his start in some scripting languages, so as to be able to relate with their "Rust is hard" experience?
Folks always say this but never really clarify why this should be the case.
When doing low lever unsafe stuff Rust does indeed make it very clear if something is a pointer or not, and avoids weird coercions (even C doesn't do well here!)
But for normal code, does it really matter if something is a pointer or not? In C++ already with by-ref and move values and stuff you have implicit pointers in many places that don't behave as pointers.
Rust never goes the way of adding implicit allocations or implicit indirection; it only goes in the way of stripping it, which is rarely if ever a problem except when dealing with unsafe code (where Rust is more explicit on pointers anyway).