What's a reference in Rust?
jvns.ca
jvns.ca
For those wanting to know more about this, the idea is that types whose size is unknown at compile-time receive this two-word representation. I tend to refer to these as "fat pointers", which is terminology from Cyclone (though Cyclone's fat pointers serve a different purpose). More documentation on these can be found at https://doc.rust-lang.org/beta/nomicon/exotic-sizes.html#dyn... and in the section in the book on slices (terminology taken from Go, whose slices are similar though with an extra word) https://doc.rust-lang.org/book/second-edition/ch04-03-slices...
Interesting; I've always thought of Rust slices as being rather different from Go slices. A Rust slice is always used through a reference and doesn't own its data, whereas a Go slice is not generally used through a pointer and sometimes points to a heap allocated section of memory, so it's basically the union of Rust's vector and slices.
Tangentially, the inability to tell whether some value is heap allocated or not from the type is one of my main gripes when working with Go as opposed to Rust; in Rust, I can be sure that `Vec`, `String`, `Box`, `Rc`, and `Arc` are all heap allocated and that slices, arrays, `str`, `&T`, and `&mut T` are not. In Go, slices and pointers might be heap allocated--or they might not.
That's not really true, because you can easily create e.g. a `&str` out of a `String` or a `&[T]` out of a `Vec<T>`.
I think what op means is both `&str` and `&[T]` live on the stack, even if they reference slices of heap-allocated (or static) data.
IIRC small fixed-size byte arrays (`[N; u8]`) are sometimes allocated on the stack depending on their size, but they are plain bytes and not full-fledged UTF8 strings. You can convert them into `String` but that would heap-allocate them.
To expand on this: in Rust copy strinctly means there is a new owner that is tasked with deallocating that data once it goes out of scope. Move is the same but the ownership is transferred (thus the old owner is no longer responsible for deallocating anything) insted of having a new copy and an additional owner. Copy always results in a new object of the same type.
& types are always borrowing. You can't copy into a reference since references are just borrowing of data, and owners of references won't deallocate anything (since they assume that, as long as they can hold an &, the data they reference is still alive, which is true because owners can't deallocate anything if there is a borrow in place).
EDIT: I was wrong. As sibling comment says, you can convert a stack-allocated fixed-size array slice into `&str` with `str::from_utf8`.
Check the last example here: https://doc.rust-lang.org/std/str/fn.from_utf8.html#examples
I think this is referring to a syntactic difference more than an implementation difference. In Rust, a &[u8] (usually called a "slice" but maybe more technically a "slice reference") is a pointer + a length. This is basically the same as Go's []byte, which is a pointer + a length + a capacity.
Rust also sometimes uses the [u8] type (without the &). This is an "exotic" type, in that it has no fixed size. It refers to the bytes inside the slice, but it's not really a pointer to them -- it is the bytes themselves, however many of them that might be. This mostly comes up when you're dealing with generic traits like AsRef or Deref, which will put the & back in all of their method signatures.
What I wanted to emphasize was that whether you're reading bytes through a Go []byte or a Rust &[u8], the same "number of hops" is happening at runtime.
Why is this important? The whole point of GC is to not spend time debating this point, as long as performance is good enough
I'm not trying to knock Go's performance here; from a naive standpoint, GC'ing only some pointers is better than GC'ing all of them like you have in more traditional garbage-collected languages, but it's still easier to know what exactly is being heap allocated and what isn't in a language like Java precisely because you know that all objects are on the heap. From what I've seen of low-level optimizations in Go code, it relies heavily on techniques like generating flame graphs to analyze where allocations are occurring, which IMO isn't a very good workflow, whereas in Rust you could do this much more easily by just looking at the types that are used. I don't think this approach is necessarily incompatible with garbage collection; theoretically a language like Go could have separate vector and slice types like Rust does, and I think that would make these types of optimizations much easier!
(I'm also not sure why you were downvoted for asking this; it's a perfectly reasonable question)
All it takes is a basic understanding of cache architecture and of generational GC, and simple data structures.
As an aside, I don't think Java actually suffers from the specific problem that I was mentioning in my original comment, namely that it's hard to tell what's on the heap or not. I was under the impression that all objects on Java are on the heap, which makes it trivial to determine whether something is heap-allocated or not based on the type like in Rust.
Nicklaus Wirth's Oberon OS (written in Oberon), Microsoft's Singularity OS (written in a variant of C#), the Mirage Unikernel written in OCaml, these are all examples of OSs written in GC'd languages. I am not aware of performance being an issue in any of these cases. Oberon was extensively used at ETH, and the components of Mirage that I am aware of (such as their OpenSSL and DNS) are competitive in performance with their C counterparts.
[1] : https://groups.google.com/forum/#!topic/golang-nuts/KJiyv2mV...
Absolutely. It's not a GC issue, it's a design issue. Adding this kind of control would make the language harder to use.
Making the hard case easier to handle for experts makes the easy case harder to handle for everyone.
Slices should be capable of being partially deallocate so long as the backing arrays are not referenced anymore.
Then to avoid performance penalties, you need to reduce allocations to the minimum, but since Go use escape analysis to decide whether to allocate on the heap or not, you don't have full control on what is heap-allocated or not, and avoiding allocations can be quite tricky.
The difficulty is knowing when that happens, and you're probably guessing wrong (the -m gcflag will tell you).
Furthermore fitting Go's theme the escape analysis is pretty simplistic, so there are many cases where it will somewhat unexpectedly assume escape (note: link is from 1.15, some have been fixed since like the …arg one or the slice assignment): https://docs.google.com/document/d/1CxgUBPlx9iJzkz9JWkb6tIpT...
if T is Box<u32>, &T is heap allocated. Imho, it's more accurate to say that `&` and `&mut` doesn't cause heap allocation. </nitpick>
let b = Box::new(5);
let r = &*b;
Here, r is on the stack (or in a register), but is pointing to something on the heap, so "&T is heap allocated" feels wrong.What's the difference between `r` and `b` here ? (no pun intended)
`b` cannot be moved, mutated, or deallocated until `r` is gone. When `b` goes out of scope the heap value it points to will be automatically deallocated, unless `r` still exists somewhere (saved off in a struct for example), in which case the program won't compile.
My question was, is it wrong to say that `b` isn't heap allocated either since : «Here, b is on the stack (or in a register), but is pointing to something on the heap.»
My knowledge of C made learning Rust so much harder for me. It's really hard to stop thinking in pointers. While Rust's references are technically implemented as pointers, for the purpose of "fighting with the borrow checker" it makes more sense to think of them as read/write locks for regions of memory.
I work at GitHub and I’ve been telling people that for the future of open source we really ought to be looking at the Rust community, both the amount of automation they have and also their general communication style.
I need a good reference on the right way to un-learn certain C concepts to make learning Rust concepts easier.
> These 3 types all have equivalent reference types (again: a reference is a pointer to memory in an unknown place): &[T] for Vec<T>, &str for String, and &T for Box<T>.
This seems to accidentally imply that these reference types are for things on the heap. i.e., that &T is borrowed equivalent to Box<T> which is not true. All three of these reference types can point to memory not on the heap. The former two 'usually' don't, while the latter will vary wildly depending on the application.
Or did you mean &T usually points to things on the heap, in which case I should just say it very very commonly points to stack allocated things as well.
Really? I would say that in my typical Rust code &[T] is created from a heap-allocated array >90% of the time. Most functions that do not require ownership of an argument will use &[T] and not &Vec<T> (or perhaps S: AsRef<[T]>), since &[T] works for stack and heap memory and &Vec<T> is automatically converted to &[T] through Deref coercion.
E.g.:
fn main() {
let v = vec![1, 2, 3, 4, 5];
blah(&v);
}
fn blah<T>(s: &[T]) {
println!("{}", s.len());
}
(The same is true for &str.)For &str you have to remember that every string literal in your program is one. When you do `some_String.starts_with("/mnt")`, `println!("hi there {}", name)`, etc you are using a new &str. I suspect most programs use more static strings than dynamic Strings (particularly since Rust isn't heavily used in GUIs yet).
> ...
> When the function blah returns, x goes out of scope, and we need to figure out what to do with its my_cool_pointer member. But how can Rust know what kind of reference my_cool_pointer is? Is it on the heap?
> ...
> If we knew that my_cool_pointer was allocated on the heap, then we would know what to do when it goes out of scope: free it!
The way this is written kind of seems to suggest that Rust will sometimes free heap memory when a reference to that memory goes out of scope, which I think is misleading.
As I understand it, this is not the case, and the point is just that Rust needs to be able to prove that nothing else freed the referenced heap memory at any point where the reference may be used.
The trick is that if it is (possibly) returned from a function, it is moved instead of dropped.
It's also important to distinguish Box from Rc. Both are heap values but have very different behavior.
The word "reference" is overloaded, it can be used to mean "anything pointery that's guaranteed to exist" too. Box<T> in this context is a reference.
The post does kind of dance between definitions of "reference" a bit, but I think that's intentional.
With that said, I'd like to add some advice by spring-boarding off a part of the post.
> Converting from a Vec<T> to a &[T] is really easy – you just run vec.as_ref(). The reason you can do this conversion is that you’re just “forgetting” that that variable is allocated on the heap and saying “who cares, this is just a reference”. String and Box<T> also has an .as_ref() method that convert to the reference version of those types in the same way.
While on the surface this is absolutely correct, there is a subtle point missing here: as_ref on Vec/String/Box is implemented as part of the AsRef[1] trait, which is _intended_ for use in generic programming. Aside from intent, practically speaking, using as_ref in a non-generic context can often be somewhat unergonomic, since depending on how you use it, it might require a type annotation (because it's generic!).
Where AsRef is useful is in making the types of parameters to functions a bit more liberal. One particularly convenient place where it's used in the standard library is for defining functions that accept file paths. For example, the type signature of the function that opens a file is[2]:
fn open<P: AsRef<Path>>(path: P) -> Result<File>
Basically, this function says that it accepts a parameter `path` with a type `P` that can be infallibly converted into a `Path`. Why is that convenient? Because lots of useful types implement `AsRef<Path>`. They include OsStr, Cow<'a, OsStr>, OsString, str, String, PathBuf, and of course, Path itself. This is what let's you write `File::open("foo/bar")`. Without the generic `AsRef<Path>` constraint, the signature would look like this: fn open(path: &Path) -> Result<File>
Which would mean that you'd need to write something like `File::open(Path::new("foo/bar"))` instead.So what's the alternative to using `as_ref` if I'm here poo-pooing it? In my experience, the typical thing to do here is to rely on something called deref. That is, if `s` is a `String` then `{STAR}s` is a `str` and `&{STAR}s` is a `&str`. In many cases, the explicit dereference (so that's `&s` instead of `&{STAR}s`) can be elided and the compiler will "auto-deref" for you. For example, given a function like the following
fn repeat(string: &str, count: u64) -> String
and a string `s` with type `String`, then repeat(&s, 5)
will "just work." If you prefer the explicit, then I think the recommendation is to use type specific conversion methods. For `Vec<T>`, `as_slice` will give you a `&[T]`. For `String`, `as_str` will give you a `&str`.OK, that's enough for now! This rabbit hole goes deeper, but I'll stop here. :)
> One question I have (that I think I will just resolve by getting more Rust experience!) is – when I write a Rust struct, how often will I be using lifetimes vs making the struct own all its own data?
If I were forced to give a pithy answer to this question, then I think I would say (predominantly from the perspective of a library writer): "It's a healthy mix, but if I don't care about performance for $reasons, I can usually ignore lifetimes in the types I define."
[1] - https://doc.rust-lang.org/std/convert/trait.AsRef.html
[2] - https://doc.rust-lang.org/std/fs/struct.File.html#method.ope...
Can't one achieve same with enum??
enum Path {
FromString(String),
FromStr(&str),
FromOsStr(OsStr),
// ...
}
fn open(path: Path) ... {
}
And then as the caller you'd have to do things like: let file = open(Path::FromStr("/foo/bar"));
It's not particularly nice to read, and you also have the overhead of creating and throwing away the enum instance.Not true, zero-sized structs are quite useful too. They can be used to fulfill traits, indicate certain errors (often in enums), etc.
A couple quick examples from the stdlib:
https://github.com/rust-lang/rust/blob/1.22.1/src/libstd/syn...
https://github.com/rust-lang/rust/blob/1.22.1/src/libstd/col...
I recommend watching this excellent rustconf 2017 talk for more information; it heavily features information on how zero-sized types can be used: https://www.youtube.com/watch?v=wxPehGkoNOw
It's sufficient and actually a lot nicer to simply state your point: e.g. "Zero-sized structs are quite useful too."
That's not true. In Java pointers can very well be allocated on the stack, but the objects that they point to will be on the heap
Can't you simply use the docs? When I code in Rust I generally have the docs opened: https://doc.rust-lang.org/std/vec/struct.Vec.html (or more like likely the locally installed version).
Most crates have documentation available as well (generally linked directly from their entry on crates.io) and if it's not online for some reason you can just run "cargo doc" to generate it locally. Randomly taking the "image" crate as an example: https://docs.rs/image/0.17.0/image/
Beats grepping header files IMO.
who greps header files in 2017 (or even 2010) ? just fuzzy search a few characters that more or less looks like what you want in your IDE's search box.
I miss being able to fuzzy search sometimes, but I keep coming back to vim. IDEs just don't cut it for me. They are too slow (Visual Studio 2017 on my desktop from 2011 is unbearable for even starting a new project). And most things I really need to do - in vim they are a few memorized keypresses or a plain shell command in a Makefile away, while in IDEs I have to dig through wizards which really brings me out of the zone.
Not relying on API search much has the huge advantage of not relying on external APIs, which leads to good modularization. As a general rule, a module shouldn't call into other modules much.
And by the way it's the same for OOP: OOP has the advantage of supporting IDE member/method autocomplete (noun first syntax), but it's just the wrong mindset for me and leads to really broken architectures.
When writing Rust, you'll likely use the standard library a lot; this rule might not be as applicable as in other languages/environments.
It seems like this information is all collected together for the docs, for example.
Look at the page for std::vec::Vec, for example (https://doc.rust-lang.org/std/vec/struct.Vec.html).
Here, you have sections for: 'Methods', 'Methods from Deref<Target=[T]>' and 'Trait Implementations', and then it seems that if you look through all these sections, you can see everything that can be called on this type, highlighted in the same light brown colour.
It would be quite nice to get an alphabetically ordered list of just these method names, also..
Like C++, you can also (ab)use intellisense to find a lot of them as well. I should hack more on Visual Rust to improve the situation there...
Trait implementations may bring in other methods and may be listed elsewhere, but C++ doesn't help with this either (C++ doesn't have traits but there are common patterns that provide similar functionality)
Most folks use the autogenerated docs (cargo doc), which list all the methods. But also when reading code it's not hard to grep for impls.
what makes them more first-class than C++ references ? eg in C++ given a type T, you can use `std::add_lvalue_reference<T>`, `std::remove_reference<T>`, overload on references, check if a type is a reference to another...
C++ references on the other hand are more like modificators of a type, eg you can have a T or a T&, but having a (T&)& does not make sense. (Outside of templates, where it gets folded down to a T&.)
This makes me feel hopeless, as I'm only about to start using Rust in my hobby projects after reading the essential book chapters. I hope it's just excessive humility on her part ? At the same time, I'm excited because if I commit myself to mastering such a language it can make me stand out. I still have an opportunity to be an early adopter, and have a head start in a promising new language.
The operative term there being "a few hundred lines", not "4 years".
A few hundred lines is hardly a couple of hours work.
This means she's written bits of Rust on and off over the course of four years and never sat down with it, basically.
Nothing to worry about.
Good luck!