Should small Rust structs be passed by-copy or by-borrow?
forrestthewoods.com
forrestthewoods.com
If you were to discuss the headline issue, I guess you could delve deep into calling conventions, register pressure, various instruction set extensions that compilers can abuse to carry the data.. or you take the compiler optimizations as a hint that concerning yourself with this is mostly pointless, as long as you don't start passing huge structures by-copy.
Trying to guide compilers with hints is a futile effort and most of the time compilers would be expending best effort to begin with.
I think it would be better for all traits to implement their methods on `self` and then the caller can choose whether the type that implements the trait is `i64` or `&i64` (instead of the trait being prescriptive about the type, which largely defeats the purpose of traits to begin with).
That said, I'm not very experienced in Rust, so I expect someone can correct me about why this is a Very Bad Idea.
There might be a way around this limitation, but it makes figuring out who is supposed to free a resource hard to do.
The rust solution is to have different traits for different levels of ownership. For example, if I want to iterate over a vector of strings, I can get a Iterator<&str> without destroying the vector. If I'm willing to destroy it, I can get an Iterator<String>.
Or, you can actually implement a trait on a reference type[0] to convert a `self` argument to effectively `&self`. This is the recommended way to implement some conversions, using AsRef and/or Into[1].
[0] https://play.rust-lang.org/?version=stable&mode=debug&editio... [1] https://doc.rust-lang.org/std/convert/trait.AsRef.html
That's exactly what the author did.
In the below example Rust switches to using a pointer when you might think it's doing pass by copy/move. This works because in Rust the moved value can never be referenced after calling the function, so the compiler can just pass a pointer and clean up the value after the function has returned.
https://www.reddit.com/r/rust/comments/3g30fw/how_efficient_...
Isn't this a negative for Rust for when you actually do care about such micro-optimizations? It feels as if you not only need to know the language itself but also how the particular version of the compiler you are using has decided to interpret the language and apply optimizations - essentially having to know how the magic trick is performed.
AFAIK this is why Free Pascal added constref in addition to const (which existed since the 90s in Delphi) - the latter doesn't guarantee pass-by-reference (even if in most cases it will do that) so if you care about that you need to know how the compiler you are using will treat it (thus needing to know more than just the language) but constref does exactly that. Now granted this was mainly done for interfacing with external code written in other langauges than micro-optimizations, but it still applies that if you cared about what the compiler will produce you had to rely on "magic knowledge" before constref was introduced.
That's also true with C/C++, isn't it? Optimizations should be different if you compile a C program with GCC or CLANG, or even with GCC x or y versions (but I'm not a compiled language programmer, I could be wrong)
I have the misfortune of working somewhere where I generally have to use older versions of compilers for production. I've found two compiler bugs just in the last two years since I started my current project. I'm usually the last person to blame the compiler, but in this case I can be pretty confident because the bugs I found were fixed in later versions of the compiler.
If you need concise pointer control, use the unsafe trapdoor built into rust (like Rust's Vec). If you need very specialised code use assembly, not by abusing language side-effects like Duff's Device. If you need fast number crunching use the specialised library for it, not by writing a magic for loop with just the right amount of statements in it.
I admit that sometimes you need to micro-optimise, but I haven't come across this case much in my Rust code. When I do I usually hinder more then help it.
There is always that group of people who are trying to do something for which no specialized library exists.
The actual generated code depends on the optimizer, while preserving semantics.
1. a large sample size
2. a single thread
3. high-precision monotonic timestamps on an unloaded system
4. mean/std dev from more than several runs
Probably not... and too much to ask.
6. for similar reasons, run them in a random order each time
This is a very obscure topic in general, and almost requires reading the C++ standard to understand it.
You can cut down on a lot of that by implementing basic operations (Add, Mul, etc.) for both value and reference types.
impl Add<Vector3> for Vector3 { … }
impl Add<&Vector3> for Vector3 { … }
impl Add<Vector3> for &Vector3 { … }
impl Add<&Vector3> for &Vector3 { … }
(The first three can delegate to the last one, unless there are value-specific optimizations you want to apply.)A crate with a derive macro for example?
impl<T> Add<T> for &Vector3 where T: Borrow<Vector3>https://github.com/isocpp/CppCoreGuidelines/blob/master/CppC...
Bigger than 3 words, pass-by-ref, otherwise pass-by-value.
So the only important modifier is whether the parameter is mutable or not. Obviously if mutable it's always passed-by-reference.
But yes, in this case it probably came down to some more obvious cause like a wildly different compiler.
That might be due to differences in the code, but it might just as well be due to differences in the flags that rustc and Clang pass to the backend by default. It's not even clear what optimization level they asked for.
AIUI it recommends by-copy for structs of at most 8 bytes because it thinks that's the appropriate limit for 32-bit targets, and they don't like the set of warnings issued to change too much for different targets.
With built-in style lints and Clippy, Rust is doing quite well in keeping codebases simple and idiomatic.
a) the language designers decided to cram several paradigms into it, including a heavier than usual dose of functional.
b) most people don't want to get pedantic about resource management
c) the syntax is weird for people coming from the C language family
Rust is actually really close to JavaScript in that regard, with its mix of OOP-without-inheritance and the whole collection of functional method on Arrays and iterators (map, reduce, etc.). And while People have a ton of complain about JavaScript, none of them is «the JavaScript language is too complex».
> b) most people don't want to get pedantic about resource management
But Rust is explicitly aimed at people who do! You can't use Python or JavaScript to write a web browser, a kernel or anything where performance is paramount.
> c) the syntax is weird for people coming from the C language family
What ? Do you mean “ not coming from the C language family”? Because its syntax is clearly in the C-family (semicolon, braces).
My guess is that they're talking about the syntax inspired from its ML roots, like `let/let mut`, `->` for return types, and `:` for type annotations. I personally prefer these to the C-style equivalents, but I've talked with enough people who haven't used an ML-like language that I'm not quite as surprised as I used to be about people expressing misgivings about them.
The nice thing about Rust is that it allows you to compartmentalize complexity very nicely by creating rules the compiler will enforce, whereas with C++ the programmer is also burdened with knowing them.
I’m not sure if I’d say that it’s more complicated than Python, since Python’s flexibility can in practice lead to creative but very hard-to-follow code.
If a significant number of the developers that come into contact with a language think it's hard, then... it's very likely hard.
They're not great, but they're a lot better than using out parameters!
I have two thoughts.
====
1. Regarding the benchmark itself, I wonder out loud if CPU caching could have a meaningful alter on the results. The article says :
> I randomly generate 4000 spheres, capsules, segments, and triangles.
It is not clear to me if this is enough to fill CPU L1 cache or not. My guess is that it is not. If all the benchmark indeed happens within the L1 cache, I wonder how a bigger workload would affect the results. Maybe it does not change anything. But maybe it does.
====
2. The author mentioned two criterion for choosing between by-copy and by-borrow : performance and ergonomics. I would add a third, somewhat related to ergonomics : let's call it semantics. When some piece of data naturally belongs somewhere and some others entities are natural "readers" of it, it might make sense to use the ownership / borrowing mechanism of Rust to naturally reflect this relationship, regardless of performance or code ergonomics.
In particular, a borrow guarantees you always work on a read-only, up to date value of the data. Meanwhile, after a pass by-copy the life cycles of the data seen by the caller and the data seen by the callee become independent of each other. This can have consequences, particularly in a multi-threaded environment.
I agree with your point 2) in general. first try for the closest mapping between language semantics and program semantics. come back and do weird stuff later after profiling.
fn dot_product(a: &Vector3, b: &Vector3) -> float {
a.x*b.x + a.y*b.y + a.z*b.z
}
fn do_math(p1: &Vector3, p2: &Vector3, d1: &Vector3, d2: &Vector3, s: f32, t: f32) -> f32 {
let a = p1 + &(&d1*s);
let b = p2 + &(&d2*t);
let result = dot_product(&(&b - &a), &(&b - &a));
}
Namely, how are you going to multiply a struct like d1 by an f32? Rust has deref coercions also, so there's plenty of times you don't even need to put a &.Also, I think the idiomatic thing to do would be to implement 'Add' for Vector3.
https://en.wikipedia.org/wiki/Scalar_multiplication
Here's a Rust Playground with the vector math impls he's probably assuming (dummied): https://play.rust-lang.org/?version=stable&mode=debug&editio...
You can get rid of the superfluous & symbols by implementing for &Vector3 also, so again, the ergonomics section doesn't really resonate with me.
Anyway good article.
That feature was in an early Modula compiler.
Surely C or C++ optimizers won't do this?
When they can prove it doesn't change the semantics on ways that violate the standard I think you will find that they do. Rust doesn't really have it's own optimizer after all, just a repurposed C/C++ one.
The problem is, "moving" a C++ object means executing an arbitrary function, which at the very least needs to put the object to a "do nothing on destruction" state. The moved-from object still exists and has to have its destructor run (which is also an arbitrary function). In Rust, even an unoptimised move is just a bitwise copy, and the moved-from object doesn't need a destructor run because its lifetime ends completely when it is moved from.
C and Rust compilers do the same thing - they follow a calling convention. It will tell you how to pass a struct, if it should be passed as a pointer or in registers etc.
Please don't blindly apply this line of reasoning everywhere.
For this particular micro-benchmark the results were almost identical. So for this particular piece of code one should apply whatever makes sense.
But if you are trying to generalize to other code (or worse yet, trying to come up with a rule), then you need to understand why the code is behaving the way it is. If you blindly trust your benchmarks you become blinded by their implicit assumptions.
https://www.forrestthewoods.com/blog/should-small-rust-struc...
I thought this strange on a 32bit machine so I dived in and it turned out the "rule" is supposed to be "up to 2 register sized" variables are supposed to be copied, anything bigger is supposed to be referenced.
The trick here is that due to compatibility and no knowlege of target platform and setup clippy assumes 32bit register size. So on 64 bit platforms you'll only get this warning up to 8 bytes as well for example.
I never went in to actually see if it makes any sense on that particular 32bit platform, so good to see someone taking a dive on the actual compiled code side.
My friend also explained how it's a REALLY tricky question to answer properly especially given specific embedded architectures and setups where the answer is very hard.
To answer that general title question, I would not even have considered that 'speed' might be an issue. I'd say: pass small things by value because it probably results in the most readable code, which should always be the first dimension of optimisation, I think. Also, the compiler is most probably good enough anyway, so care about your source code readability first.
So, at least the title should say 'if you care about speed a lot'.
The challenge of Rust is, how nice can we make a language without noticeably giving up any performance to C++? Pretty nice, it turns out.
C++ still has an edge on both performance and ability to code powerful libraries that are not a PITA to use. The performance margin definitely will close up. The library margin could, too, but it will require unpleasant choices.
Anyway both are faster than almost anything else, and will stay that way, for von Neumann machines. Those might get less important, soon.
I use Rust and I actually care about readability first.
Sincerely I find Rust more readable than most popular languages (Python, Javascript, C++, Java, even Swift...), unless I'm looking into some code that abuses generics and lifetimes (which is actually rare for me).
I think that's because it enforces some pretty clear patterns with its lack of struct inheritance, sum types with pattern matching, the way things are imported ("use" statements), lack of parenthesis around conditionals, and being expression oriented.
What was the third thing??
The author is making the case for special treatment of structs for a special case. For a performance gain that's probably below the margin of error. This is negating the elegance and consistency of rust for a special case. I think the author thinks too much in C++ terms and is sort of missing the point of Rust (or at least one point of rust).
The struct by-copy is mostly a good API thing in a lot of scenarios, particularly when the objects are <1 cache line.
Even more so, when you're only writing half the code and want to separate the actions within a function from any side-effects outside.
I would say that the performance difference for small structures is irrelevant to the extra hop, particularly if the copied location is the stack on the callee.
Due to the default immutability of references and the fact the compiler won't allow you to won't allow you to share data between threads (without jumping through a lot of hoops), the risk of side-effects are pretty well neutralized. Another selling point of Rust.
It probably will be.
I've been using Rust in production for a while now, and our bottlenecks are never Rust's performance. Something else is the bottleneck way before we approach anything close to Rust's peak throughput. This is especially true if you're doing anything that touches networking or interacts with other services.
let ref a = p1 + &(d1*s);
let ref b = p2 + &(d2*t);
dot_product_by_borrow(&(b - a), &(b - a))1) What makes the most sense semantically? (So that using your API correctly is the most obvious path)
2) Doesn't matter? Okay, what's the most ergonomic?
3) That code path needs to be performant? Try variations and benchmark within your overall application.
I'm not refuting/dismissing the article. Dives into how the compiler handles things is always interesting! But the programming ecosystem is going through some growing pains right now.
At the beginning of time we had the Era of Assembly; back when our machines were measured in MHz. Code had to be performant above all else; that was the only rule.
During the Era of Moore, performance exploded. As a counter-reaction to the Era of Assembly, programmers began chanting the mantra "No Premature Optimization!" This led to the creation of easy/lazy languages like Python, where code was an art and performance wasn't even an afterthought.
Now we're in the Era of Types. While Python and Javascript were running rampant without types (quack quack), languages like Haskell were evolving in the shadows with intricate and expressive type systems. The fruits of those labors are sprouting in the form of Rust. The mantra of "No Premature Optimization!" isn't enough. We no longer need to choose between code that is easy to write and code that is performant. With a good type system we can be more explicit with the semantics of our code and APIs. This makes our code easier to use, more ergonomic, and gives the compiler more information that it can leverage to optimize our creations into assembly machines that would make C64 developers nod approvingly (though knowing secretly that they could always do better).
The growing pain is this transition from a mindset of just "No Premature Optimization" where the focus is simply on writing "easy code" in stark reaction to the hyperoptimization of the Era of Assembly, to a modern mindset where we should write code with intentionally designed types and semantics. Hence my more complicated list of three rules instead of one.
Side Note: And of course, as some comments have pointed out, step number 3 of my guideline is fraught with peril because the compiler's behavior, and the behavior of the CPU, change like the direction of the wind. Thus the thrust of my comment is emphasizing the use of our new typing systems. If your code is typed correctly and your intentions clear (per step #1), the compiler will do the right thing on average at least.
Side Note 2: None of those guidelines apply to languages of the older eras; they aren't expressive enough to communicate with the compiler and hence you're left in the old miasma of untyped languages where optimization is near hopeless, or languages with type systems so narrow that you're battling day-to-day to build APIs with a semblance of usability.
The case in point, subroutine calls, the call itself is far costlier than the argument copying. So in this case, data moving is almost completely irrelevant.
If the compiler is smart enough to inline, it should be able to notice that the way arguments are passed make no difference and use the most efficient way regardless of the signature. Do I overestimate the abilities of compilers, or maybe there is some side effect I am not aware of (concurrency?).
Obviously, it only applies when inlining.
Regardless, I want a library author to help me fall into the pit of performance success, not expect me to do it all for myself.
fn foo<T>(frob: T) where T: Borrow<Frob> {
// Do something with frob.borrow()
}
Due to monomorpization, foo is specialized for moves/clones when called as foo(frob) where frob is a value type or for borrows when called as foo(&frob). Similarly, you can use Clone, ToOwned or even Into for going in the opposite direction.Whether such functions are good taste in general is debatable ;).