Should small Rust structs be passed by-copy or by-borrow? (2019)
forrestthewoods.com
forrestthewoods.com
Unless you are gonna benchmark something, for details like this you should pretty much always just trust the damn compiler and write the code in the most maintainable way.
This comes up in code review a LOT at my work:
- "you can write this simpler with XYZ"
- "but that will be slower because it's a copy/a function call/an indirect branch/a channel send/a shared memory access/some other combination of assumptions about what the compiler will generate and what is slow on a CPU"
I always ask them to either prove it or write the simple thing. If the code in question isn't hot enough to bother benchmarking it, the performance benefits probably aren't worth it _even if they exist_.
How are you using the word simpler? Because to me that implies a combination of more obvious and number of lines of code. Something that a benchmark shouldn't be involved in.
For example asking someone to delete 10 lines of code and instead use go's ` net.SplitHostPort` would be an example of "simpler".
Good programmers play code golf, great programmers write readable and maintainable code.
Your example seems reasonable but programmers also like to act like the smartest one in the room. I often come across tricky and borderline obfuscated code because somebody wanted to look clever. This is a logistical nightmare.
My general rule is that if you need fewer lines of code to implement your logic with a simple for loop, you probably should.
Javascript (v8, last I checked) is the opposite. Simple for loops almost always outperform anything else.
Most of the time you should just write whatever's clear/convenient but sometimes it's worth trying both and scrutinizing godbolt.
It’s also much easier to stick to for() loops in javascript as you code than it is to rewrite everything later when tuning for performance. If that’s something you expect to need to do.
Even if it's not smart enough to do that today, it could implement this optimization in the future. This could work even without inlining, since the Rust calling convention is unstable, and an optimization based on type size could be incorporated into it.
> I always ask them to either prove it or write the simple thing. If the code in question isn't hot enough to bother benchmarking it, the performance benefits probably aren't worth it _even if they exist_.
One of my philosophies is that death by a thousand cuts is fine, but death by ten thousand cuts isn’t. A team of 10 engineers can probably fix most of a thousand cuts in two or three months. But if you have ten thousand cuts you’re probably doomed. And those don’t show up cleanly in a flame graph.
Now for some context my background is video games. Which means the team knows they need to hit an aggressive performance bar. This isn’t true for many projects. shared_ptr is a canonical example of death by ten thousand cuts.
That said, I strongly agree with the principle of “just do the simple thing”. However I think it’s important to have “sane defaults”. A project can easily have a thousand or ten thousand papercuts that kill performance. But you can’t microbench every tiny decision. And microbenches are only a vague approximation of what actually matters.
I’m also wary of “the compiler will make it fast”. Because that’s true… until it’s not! Although these days you don’t have any choice but to lean heavily on the compiler and “trust but verify”.
No one wants a super complex solution if it’s not needed. However I am very amenable to “do a slightly more complex thing if you know it’s correct and we can never think about this ever again”. It’s much easier to do the fast thing upfront than for someone else to try and speed it up in two years when we’re doing a papercut pass.
I am reminded of the lovely nanosecond/microsecond talk by Grace Hopper. If your code does a little bit of setup and then spends all of its time in a single hotspot, fine. But if your code is full of microsecond-suboptimal speed bumps, you can probably hide your hotspot altogether. And a flat-ish flame graph looks fine: nothing stands out as a problem!
It's valuable to do micro-benchmarks, not just to hone your optimization skills, but to learn optimal patterns in your language of choice. Then, when you're "in the zone" and laying down new code, you just do the optimal thing reflexively. Or, when you're reviewing or rewriting something, those micro-hotspots jump out and grab your attention.
There's a reason that ancient software running on ancient hardware is way more responsive & snappy than what we have today. Laziness.
Laziness in terms of using entirely inappropriate algorithms, sure.
Laziness in not microbenchmarking minutia? It shouldn't be. There's a limit on how much that can hurt you. I would say much less than a factor of ten, but let's go with 10x just for argument's sake. If you have a CPU that's 500x faster, and use easy code that's 10x slower, you're doing just fine. This is not the problem with modern unresponsiveness.
People have lionized Knuth's quote about premature optimization, and used that to ignore performance issues across the board. Since the early '00s, we have not seen a 500x improvement in CPU speed. It's less than 2x on frequency, and let's say 8x on core-count for most users (which doesn't help your single-core lazy programmer). In my experience, programmers will make projections based on a 500x-faster processors that will never arrive, because it's easier than honing their skills and keeping them sharp. And even if these magical THz-frequency chips arrive, if you have three layers of 10x slowdowns, you're back down to GHz.
Automerge’s design assumed this stuff would always be slow, so they had this whole frontend / backend code split so they could put the expensive operations on worker threads. Good optimizations in the right places make all that complexity unnecessary. The new automerge is shaping up to be simpler as well as faster.
And neither of those is a microbenchmark thing, which is kind of my point. I'm surprised language would hurt that much, but that's enough to break things on its own without any layering.
> Since the early '00s, we have not seen a 500x improvement in CPU speed. It's less than 2x on frequency, and let's say 8x on core-count for most users (which doesn't help your single-core lazy programmer).
I don't think people are talking about 2004 when they talk about the responsiveness of ancient software on ancient hardware. I interpret that as more like an Apple II. But instructions per clock have also gone up a lot since the pentium 4 days, and having more than one core in your CPU has a huge impact even for single-threaded programs.
The point I'm making here is that every line matters -- not just the hotspots. If you're suprised that language can have that much impact, perhaps it's time to learn a bit about performance issues that you're being dismissive of?
> I don't think people are talking about 2004 when they talk about the responsiveness of ancient software on ancient hardware.
No, I was responding to your mention of a 500x improvement in hardware. That pipedream ended in the early 00s, and people still talk like Moore's law will absolve their inattentive coding practice. And that felt fine in the decades we went from kHz to GHz, but it's unacceptable today.
That depends on how it would have performed if you only transformed the hottest 10% into C.
But I was mainly responding to the idea that micro-optimizations are needed to keep general software snappy, and I don't think they are. If one language is that much faster, that's not micro-optimization.
> No, I was responding to your mention of a 500x improvement in hardware.
What do you mean "No"? I was talking about current hardware being 500x faster than 1988 hardware, which it is. If that's not what you meant by "ancient software on ancient hardware", fine, but that's what my 500x was talking about.
> people still talk like Moore's law will absolve their inattentive coding practice
I'm not trying to excuse inattentive coding. I'm trying to say certain kinds of attention are important and others aren't.
If your program is still slow, that would also indicate that everything is a problem, ie. the ten thousand cuts. Start optimizing at some obvious spots and then see what happens.
> shared_ptr is a canonical example of death by ten thousand cuts
Why does that count as ten thousand cuts rather than one cut? That doesn't sound intractable to fix if you have months.
Very few programs have "hot loops" as small, contained things. "hot loops" is at this point as much of a bad trope as "but premature optimizations!" is.
> Why does that count as ten thousand cuts rather than one cut? That doesn't sound intractable to fix if you have months.
Why spend months re-writing code you could have just written correctly the first time if a single person had spent an hour messing around with benchmarks to figure out the proper guidance?
If you say so. I can pretty easily and meaningfully divide my current program into "happens once per frame or less" and "happens hundreds or more times in a row".
> Why spend months re-writing code you could have just written correctly the first time if a single person had spent an hour messing around with benchmarks to figure out the proper guidance?
If they could have figured it out that easily I understand even less how this example works.
Let me put it this way: If a benchmark can change huge swaths of significantly performance important code, then it's "hot enough to bother" by orders of magnitude. I thought we were talking about microbenchmarks for individual cuts!
And actually I still disagree - e.g. I once took over a DMA management firmware and the TL told me "we are really trying to avoid DBATC so we take care to write every line efficiently". But the thing was that once you have a holistic understanding of the systems performance you tend to find _only a small fraction of the code ever affects the metrics you care about_!
E.g. in that case the CPU was so rarely the bottleneck that it really didn't matter, we could have rewritten half the code in Python (if we'd had the memory) without hurting latency or throughput.
Admittedly I can see how games or like JS engines might be a kinda special case here, where the OVERALL compute bandwidth begins to become a concern (almost like an HPC system) and maybe then every line really does count.
Please, don’t optimize unless you have reasons to do so and numbers backing that up.
But in practice what we see overwhelmingly is that people want to do this stuff but they aren't measuring, because measuring is boring whereas making the code more complicated to show off how much you think you know about optimisation is easy. Knock it off.
You edited your comment, so I guess I will too: Rust doesn't actually have Linear types. Linear types ("must use or compile error") would be tricky to provide, Aria blogged about it back in the day. So that's definitely going to be a problem with your refactoring.
But in my field (Fintech), performance really does matter. Doing the simple, slow thing is just lazy and won't make it through review.
That said, I think asking a developer to write everything they do twice so that they can A/B test is overboard. You can come back and really aggressively optimize later, but I think the "default" should be the fast thing, rather than the slow & easy thing.
You get just as much benefit by assigning a performance refactor to an engineer when needed vs literally halving or worse the whole teams productivity.
Edit: also optimizations often interact with each other and you can't always benchmark all possible combinations, so sometime you do have to rely on experience or reasoning from first principles to decide if an optimisation is worth it or not.
Having real benchmarking data is eye opening. It was shocking to me how much the wasm bundle size increased when I added serialisation code. The time to serialise / deserialise a big chunk of data for me is 0.5ms - so fast that it’s not worth more microoptimizations. Lots of changes I think will make the code slower have no impact whatsoever on performance. And my instincts are so often wrong. About 50% of microoptimizations I try either have no effect or make the code slightly slower. And it’s quite common for changes that shouldn’t change performance at all to cause significant performance regressions for unexpected reasons.
I’ve also learned how important “short circuit” cases can be for performance. Adding a single early check for the trivial case in a function can sometimes improve end to end performance by 15-20%, which in already well tuned code is massive.
Performance work is really fun. But if you do performance tuning without measuring the results, you’re driving blindfolded. You’re as likely to make your code worse as you are to make it better. Add benchmarks.
This is only blanket advice for designing traits, because as the trait author you don't know what concrete type the downstream user is going to want to use, and taking `&self` in that circumstance is the choice that is friendliest to both Copy and non-Copy types.
If you're just writing a non-generic function and you do know what concrete types you're using, the flowchart is pretty simple:
1. If the type is not Copy, then pass by-ref if you just need to read the value, pass by-mutable-ref if you just need to mutate the value, and pass by-value if you want to consume the value.
2. If the type is Copy, then pass by-value, but if your type is really big or if benchmarking has determined that this is a critical code path then pass by-ref.
Generally I find that less bugs get introduced when using copy instead of pass by reference, but I’m sure others have the opposite opinion.
And the objects have to get surprisingly large before passing by reference really makes a difference.
I got very little experience in rust though, so there might be a way (I'm just not aware of) to circumvent this check
https://doc.rust-lang.org/reference/interior-mutability.html
Usually (read: near universally) something that provides interior mutability will implement a runtime check to protect you (eg, a lock), but you could make a footgun if you were cavalier about it. And recent I misused a framework (I kept a value returned by the framework around that should've been dropped) without realizing it was holding locks internal to the framework, and ended up with a tricky to diagnose deadlock (though only tricky before I hadn't read the documentation as closely as I should've).
I really dislike these takes. I see engineers optimize these cases and then go ahead and make two separate SQL queries that could be one, ruining any false optimization gains they got by lord know how many times.
Yeah, you can loop over that 100 element list twice doing basic computation if you want, it's not going to make a difference for many engineering workloads, but could make a big difference in readability.
Even if they do, they also need to make a case that in this specific case, performance matters enough to pessimize code simplicity and maintainability.
Also, if performance is that critical, it's imperative to benchmark again after each compiler release to guard against codegen regressions. And benchmark after changing this piece of code. Otherwise, we can say that performance doesn't really matter.
I've seen this happen a lot with JavaScript, particularly in the last 5-10 years as JS engines have developed increasingly sophisticated approaches to performance. Today's optimization can be tomorrow's de-optimization. Even given an unchanging landscape of compiler/interpreter, tightly-optimized code can become de-optimized when updated and extended, as compared to maintainable code that may not suffer much performance degradation upon extension.
Good thing modern register allocation by compilers makes this irrelevant.
This is cumulative thou, 1 may not make much difference. 100,000 of those will. There is that story of chrome slows down that was does discovered to be in strings each one of which is too little to make a difference
Rust - Windows - By-Copy: 14124, By-Borrow: 8150
C++ - Windows MS Compiler - By-Copy: 12160, By-Ref: 11423
C++ - Windows LLVM 15 - By-Copy: 4397, By-Ref: 4396 Delta: -0.0227428%
So it appears that C++ - Windows LLVM 15 beats Rust by large margin.Of course, if your struct is truly enormous, you may want to break this rule to avoid large copies. But in that case you probably want to Box<T> the struct anyway.
Of course, if your struct contains something that can't be copied--like a Vec<T>--you'll have to decide whether to clone the whole struct (and thus the vector in it), pass the struct by-borrow, or find some other solution.
Sure, but it's worth noting that references in Rust do not exist merely to avoid passing by-value. They also exist to make it easier to deal with Rust's ownership semantics: they let you pass things to a function without also requiring the function to "pass back" those things as returned values. In other words, references let you do `fn foo(x: &Bar)` rather than `fn foo(x: Bar) -> Bar`. This is a unique and interesting consequence of languages with by-default move semantics.
(The other direction is trickier, since by-ref implies the desire to observe aliasing, even though that's usually not expected in practice - but the compiler cannot tell.)
> Then you would actually be testing by-borrow be by-copy instead of how good rust is at optimizing.
I don’t think the question is actually: “what is faster in practice, a by-copy method call or a by-value method call”, I think the question is: “as an implementer, which semantics should I choose when I’m writing my function”.
For the second question: “Rust is usually pretty good at aggressively inlining, so… if you’re willing to trust Rust’s compiler, you’re often okay going with by-copy implementations, but you should keep an eye on it”. Whereas, as you note, for the first question it’s not an answer.
But, I do think if someone was going to put more work into it I’d be very curious what the answer to the first question is. If I’m choosing to implement with by-copy semantics and trusting the Rust compiler to hopefully inline things for me, I’d like to know the implications in the cases when it doesn’t.
The root question is indeed “what semantics should I use”. And the answer I came up with was “the compiler does a lot of magic so by-copy seems pretty good”. I agree with the previous commenter this is not a satisfying conclusion!
My experience with Rust is that it requires a moderate amount of trust in the compiler. Iterator code is another example where the compiler should produce near optimal code. Emphasis on should!
Your tests on my PC:
Rust - By-Copy: 14124, By-Borrow: 8150
C++ - By-Copy: 12160, By-Ref: 11423
P.S. Just built it using LLVM under CLion IDE and the results are: G:\temp\cpp\rust-cpp-bench\cpp\cmake\cmake-build-
release\fts_cmake_cpp_bench.exe
Totals:
Overlaps: 220384338
By-Copy: 4397
By-Ref: 4396 Delta: -0.0227428%
Process finished with exit code 0I believe that the Rust compiler at least does exactly that. Large structs will be passed by reference under the hood even if it passed by value in the code. I suspect C++ compilers do the same, although I'm not sure about that.
G:\temp\cpp\rust-cpp-bench\cpp\cmake\cmake-build-
release\fts_cmake_cpp_bench.exe
Totals:
Overlaps: 220384338
By-Copy: 4397
By-Ref: 4396 Delta: -0.0227428%
Process finished with exit code 0Were the other benchmarks run in debug mode / with optimizations turned off or something like that? What compiler & flags are you using?
Why would I do something like that? Of course all builds are release mode, optimize for speed.
Rust - Windows - By-Copy: 14124, By-Borrow: 8150
C++ - Windows MS Compiler - By-Copy: 12160, By-Ref: 11423
C++ - Windows LLVM 15 - By-Copy: 4397, By-Ref: 4396
>"Why is performance so much better in this case?"Not sure and not in a mood to investigate. I do know if cache locality and branch prediction stars line up properly the performance difference can be staggering. Maybe LLVM has accomplished something nice in this department.
C++ MSVC: By-Copy: 12,077 By-Ref: 11,901
C++ Clang: By-Copy: 5,020 By-Ref: 5,029
Rust: By-Copy: 3,173 By-Borrow: 3,148
All on Windows, and on the same i7-8700k desktop I used for the original post in 2019.
Your Rust numbers are particularly curious. Maybe run `rustup update` and try again?
Rust - By-Copy: 2685, By-Borrow: 2694
C++ - Windows MS Compiler - By-Copy: 12160, By-Ref: 11423
C++ - Windows LLVM 15 - By-Copy: 4397, By-Ref: 4396
My CPU is AMD Ryzen 5950X so it seems like Rust kicks the shit out of C++ in this case. I am going to try LLVM 16 and GCC tomorrow.Happy New Year
rust - By-Copy: 2683, By-Borrow: 2697
c++ - By-Copy: 2577, By-Ref: 2600
Either is way faster than the result from the original post and no rust does non win this "competition". C++ is a bit faster (see also the results from other poster above) but not by much. rustc 1.58 (LLVM 13): By-Copy: 10804, By-Borrow: 7198
rustc 1.64 (LLVM 14): By-Copy: 7385, By-Borrow: 7328
rustc 1.66 (LLVM 15): By-Copy: 2667, By-Borrow: 2777
clang++ (LLVM 14): By-Copy: 2439, By-Ref: 2589
clang++ (LLVM 15): By-Copy: 2473, By-Ref: 2556
When compared with the same LLVM version clang much better utilised LLVM-14 than rustc - C++ is 3x faster. With LLVM-15 they are much closer. clang++ -O3 -mavx2 fts_cpp_copy_bench.cpp -o test.exe
this did it. The final score is: rust - By-Copy: 2683, By-Borrow: 2697
c++ - By-Copy: 2577, By-Ref: 2600
so C++ seems a bit faster but not by muchTry this:
RUSTFLAGS="-C target-cpu=native" cargo run --release # or whatever
And: clang++ -O3 -march=native fts_cpp_copy_bench.cpp -o test.exe
arch=native will also activate SSE, MMX, and all the other goodies that modern CPUs have to offer. rust - By-Copy: 1831, By-Borrow: 1850
c++ - By-Copy: 2411, By-Ref: 2458
Interesting why Rust wins with such a large margin. Maybe will try to find out later.But also, the ABI only matters between two pieces that are separately compiled. A static binary optimized at link-time doesn't have to care.
Also Fortran has "in", "inout" and "out".
It can return multiple values, so this doesn’t matter much for value types, but it would be nice to be able to specify that a pointer arg is an out-param sometimes and enforce that it is not read from while handling allocation in the caller.
Better yet when you prohibit such arguments from aliasing (or at least make no-alias the default) - now the compiler can also implement "in out" by copying the value back and forth, if it's faster than indirection.
Herb made a proposal for proper in/out parameters for C++ in 2020 https://youtu.be/6lurOCdaj0Y
in - regular function arguments
inout - mut function arguments
out - function return
Is there any additional information that a compiler can infer from Ada’s parameter syntax?
Back in the ancient days, I worked at IBM doing benchmarking for an OS project that was never released. We were using PPC601 Sandalfoots (Sandalfeet?) as dev machines. A perennial fight was devs writing their own memcpy using dst++ = src++ loops rather than the one in the library, which was written by one of my coworkers and consisted of 3 pages of assembly that used at least 18 registers.
The simple loop was something like X cycles/byte, while the library version was P + (Q cycles/byte) but the difference was such that the crossover point was about 8 bytes. So, scraping out the simple memcpy implementations from the code was about a weekly thing for me.
At this point, we discovered that our C compiler would pass structs by value (This was the early-ish days of ANSI C and was a surprise to some of my older coworkers.) and benchmarked that.
And discovered that its copy code was worse than the simple dst++ = src++ loops. By about a factor of 4. (The simple loop would be optimized to work with word-sized ints, while the compiler was generating code that copied each byte individually.)
If you are doing something where this matters, something like VTune is very important. So is the ability to convince people who do stupid things to stop doing the stupid things.
And I would bet 9 times out of 10 it won't be the bottleneck or even make a measurable difference.
That's my approach too as a Rust newbie. Borrow by default and take ownership only when needed, for the best ergononmics.
I don’t think you ever have to write code like this. Implement your math traits in terms for both value and reference types like the standard library does.
Go down to Trait Implementations for scalar types, for instance i32 [1]
impl Add<&i32> for &i32
impl Add<&i32> for i32
impl Add<i32> for &i32
impl Add<i32> for i32
Once you do that your ergonomics should be exactly the same as with built in scalar types.
Lots of criticism of my methodology in the comments here. That’s fine. That post was more of a self nerd snipe that went way deeper than I expected.
I hoped that my post would lead to a more definitive answer from some actual experts in the field. Unfortunately that never happened, afaik. Bummer.
And that doesn't help at all if you're writing a "free function" like 3D primitive intersection functions. I suppose you could change that simple function into a generic function that takes AsDeref? Bleh.
If there's no inlining at play, I'd expect vast differences to be possible. For example, imagine a chain of 3 functions - f calls g, g calls h, where one of the arguments is a 1kB struct and the options are passing by copy or by borrowing. In this case, each stack frame will be 1kB in size in the copy case and there will be a large performance overhead as opposed to the by-reference case. One would expect simply calling the function to be similar in overhead to an uncached memory load.
Within a single crate the inlining is possible, with multiple crates it's only possible with LTO enabled (and I'm not sure how _probable_ it is that the inlining would occur).
In either case, the difference between a 32 byte and 8 byte argument in terms of overhead is likely meaningless - the sort of thing to be optimized if profiling says it's a problem as opposed to ahead of time.
Cross-crate inlining happens all the time. In order to be eligible for inlining, a function needs to have its IR included in the object's metadata. This happens automatically for every generic function (it's the only way monomorphization can work), and for non-generic functions can be enabled manually via the `#[inline]` attribute (which does not force inlining, it only makes it possible to inline at the backend's discretion).
However, as you, say, if you have LTO enabled then "cross-crate" inlining can happen regardless, since it's all just one giant compilation unit at that point.
So, what the author actually measured was the difference between llvm and msvc throughout the article. Particularly when they talked about rust being better at autovectorization than C++.
The other question I have is which style should you use when writing a library? It's obviously not possible to benchmark all the software that will call your library but you still want to consider readability, performance as well as other factors such as common convention.
The clarity of the code using a particular library is such an big (but often under-appreciated) benefit that I would heavily lean in this direction when considering interface options. My 2c.
Situations otherwise are the exception, rather than the rule, and it takes an expert to (1) recognize those situations and (2) know exactly how to write optimized code in that situation.
That's why "don't prematurely optimize" is a good rule of thumb - because it works the majority of the time, and it takes experience to know when not to apply it.
[citation needed]
This claim depends hugely on the industry you're actually working in and the problem space. Things like UIs & games basically never have a single, easy-to-fix smoking gun. The entire app is more or less a hotspot - be it interactive performance, startup performance, RAM usage, or general responsiveness.
And once you're gone down the route of "build it first, optimize it later" you're pretty much fucked when you get to the "optimize" step because now your performance mistakes are basically unfixable without a rewrite - every layer of your architecture has issues that you can't fix without drastic overhauls. It would have been much easier to do some up-front measurements, get some guidelines in place (even if they aren't perfect), and then build the app.
> The true root of all evil is unexamined dogma.
This is so absurd that it doesn't deserve a reply.
There’s no hard and fast rule here. Even if there was, optimizers still occasionally surprise seasoned native devs in both positive and negative ways.
Glad the author’s first instinct was to pull out profiling tools.
Trusting the compiler also means knowing what the compiler actually understands & handles vs. what's a library-provided abstraction that's maybe too bloated for its own good and that quickly becomes "not simple" depending on your language of choice.
Not really. I'm just hoping that it will be "fast enough", which in the vast majority of cases it is.
* Sprinkling & around everything in math expressions does make them ugly. Maybe rust needs an asBorrow or similar?
* If you inline everything then the speed is the same.
* Link time optimizations are also an easy win.
Do you mean AsRef, or do you mean magic which automatically borrows parameters and is specifically what rust does not do any more than e.g. C does?
Though you can probably get both if the by-ref version is faster (or more convenient internally): wrap the by-ref function with a by-value wrapper which is #[inline]-ed, this way the interface is by value but the actual parameter passing is byref (as the value-consuming wrapper will be inlined and essentially removed).
FWIW, the `Borrow`, `AsRef`, and `Deref` traits all exist to support different variants of this.
References may get optimized to copies where possible and sound (i.e. blittable and const), a common heuristic involves the size of a cache line (64b on most modern ISAs, including x86_64).
Using a Vector4 would have pushed the structure size beyond the 64b heuristic. You would also need to disable inlining for the measured methods.
And, for simple operations like this, you really should just look at the assembly output. If you are only generating 20ish instructions, then look at those 20 instructions rather than trying to heuristically guess what is happening.
It does sometimes matter though. One optimization I’ve seen in a few places is to box the error type, so that a result doesn’t copy the (usually empty) error by value on the stack. That actually makes a small performance difference, on the order of about 5-10%.
The cost of by-value lies in memory copies, while the cost of by-reference lies in dereferencing pointers where the values are needed, which might mean many more memory reads are needed than with by-value (depends on what you're doing). So it's just hard to tell which will do better in general -- there's no answer to that.
For a library, maybe providing by-value and by-reference interfaces should be good (except that will bloat the library). For everything else just use by-value as it has the best ergonomics.
Rust - By-Copy: 14124, By-Borrow: 8150
C++ - By-Copy: 12160, By-Ref: 11423
P.S. Just built it using LLVM under CLion IDE and the results are:
G:\temp\cpp\rust-cpp-bench\cpp\cmake\cmake-build-
release\fts_cmake_cpp_bench.exe
Totals:
Overlaps: 220384338
By-Copy: 4397
By-Ref: 4396 Delta: -0.0227428%
Process finished with exit code 0Now comes big surprise: I just built it using LLVM under CLion IDE and the results are:
G:\temp\cpp\rust-cpp-bench\cpp\cmake\cmake-build-
release\fts_cmake_cpp_bench.exe
Totals:
Overlaps: 220384338
By-Copy: 4397
By-Ref: 4396 Delta: -0.0227428%
Process finished with exit code 0But at the risk of loss of respect, I'll wait for Rust2ShinyNewLanguage to solve this.
All I know is I hope I'm smart enough to understand ShinyNewLanguage's compiler. Or maybe even build it.
I've got several projects that could use some additional Boxes of structures, and borrow instead of move, and maybe a few more complex reference counting mechanics.
Rust forced me to understand what that meant. That's good for building a better engineer.
But it's not fun to work with.
I hope the next experience is better. Sorry Rustaceans.
The benchmark made here could completely fall apart once more threads are added.
Modern computer architectures are non-uniform in terms of any kind of memory accesses. The same logical operations can have extremely varied costs depending on how the whole program flow goes.
Anyway, assuming it's not inlined I would guess pass-by-copy, maybe with an occasional exception in code with heavy register pressure.
Edit: Actually since it's a structure, the calling convention is to memory allocate it and pass a pointer, doh. So it should actually compile the same.
Depending on calling convention, the structure may be spread out into registers.
FWIW the AMD64 SysV v1.0 psABI allows structures of up to 8 members to be passed via registers. Though older revisions limit that to 2 (and it's unclear whether MS's divergent ABI allows aggregates to be splat at all.
As sad as it's unsurprising, it does not look like LLVM (linux?) has followed up, on godbolt a 2-struct passes everything via registers but a 3-struct passes everything via the stack. Maybe there's a magic flag to use the 1.0 ABI, but a quick googling didn't reveal one. ICC doesn't seem to have followed up either.
Actually reading the word abi made a little neuron light up. To what extent does rust even follow that abi?
Also, whenever you do one of these, please post the full source with it. There's no reason to leave your readers in the dark, wondering what could be going on, which is exactly what I'm doing now, because there's almost no excuse for c++ to be slower in a task than rust--it's just a matter of how much work you need to put in to make it get there.
There’s literally a section called Source Code…
But yeah I have no idea why a benchmark framework wasn't used for Rust.
Performance-wise, if you're likely to touch every element in a type anyway, err on the side of copies. They are going to have to end up in registers eventually anyway, so you might as well let the caller find out the best way to put them there.
But also, the struct is 3x32 bits, and Rust auto-implements the Copy-trait for it. It is barely larger than u64, which is the size of the reference.
But life is only simpler when Copy and Clone can be auto-implemented.
But nevertheless: I agree it would have been interesting to test with GCC as well.
Should small Rust structs be passed by-copy or by-borrow? - https://news.ycombinator.com/item?id=20798033 - Aug 2019 (107 comments)
Rust is probably better used for writing fast low-level libraries that you call from higher level languages, possibly with a garbage collector, so you don't waste time thinking about memory management while you design/write your high-level application.
What did you have in mind?
I'll take Rust shouting at me for missing "mundane details" any day of the week.
Rust uses syntax that feels familiar but means completely different things than in pretty much any other language.
For example '=' doesn't mean assign handle or copy. It by default means move.
'let' doesn't mean create a name for something. It means create physical space for something (of known size) that can be moved into or moved out of.
You don't deal with objects and values of primitive types. Instead everything in Rust is a value. When you move, you move the value. If you compare, you compare by value. If you pass something from variable into function, you move the value into the function.
And when the space where you keep the value goes out of scope value dies with it if it wasn't moved out to somewhere else.
Scope for variables (which are just named spaces for values) ends with the end of the block, but some values, created by functions and returned from them, if they are not moved into any named space, can die sooner, even in the middle of the line where they were acquired from function call.
Everything else stems from that fixed size moved value semantics. If you don't want to move the value into the function when you call it you need to pass something else instead, so you create and pass in the borrow. But you have to ensure that the value doesn't die or get moved anywhere (even inside the container you borrowed from) before borrows to it all die.
Because of this you are better off with borrows that are short lived and local. Often it's better to keep the index of an element of a Vec instead of the borrow of this element. If you must create types that contain borrows you must know that they become borrows themselves and you need to treat them exactly the same trying to limit their scope and life time.
It's hard when you come from any other language because borrows are superficialy similar to pointers or references to objects. So you try to use them as such. And crash into the compiler because they are not that. What's worse their syntax is very minimalistic which triggers intuition that they must be fast and optimal solution for many problems which they sure can be once you fully internalize their limitations but not a moment sooner.
Another thing is that values in Rust must have the fixed size. So even as simple thing as a string requires a bit of hackery. Basically in Rust the default strategy to have something of variable size is to allocate it on the heap and treat pointer to it (possibly with some other fixed sized data like length) as the fixed sized value you can move around clone and borrow.
So if you want to have semantics you know from other languages you can't just use basic Rust syntax.
You need constructs such as Box and Rc, Cell, RefCell. Make your things clonable and sometimes even copyable and avoid creating borrows whenever possible initially. When you do it Rust becomes as flexible as any other language and you can use it pretty much just as comfortably. Then the value semantics shines as you can very easily compare your data by value, order it, create operators for it, create has for it so you can keep it in HashMaps and HashSets. Then it's delightful.
My advice is when you create a long lived type just wrap it in Rc and treat this Rc as your 'object'. And avoid borrows in your types unless you have a very good performance (measured) reason to have them or you are creating something obviously dependant and usually short lived like an iterator.
One non-tree cross-link or back-link and you'll have to redesign your entire code.
You can fairly easily refactor your almost-tree code to adapt it to that additional Option wrap.
Of course you might instead opt to introduce some garbage collector crate into your project. They usually provide garbage collected Rc equivalent, which makes swapping it out very easy.
Rc's are really very useful first approach to making anything complex in Rust.
I usually have something like
struct NodeStruct {
my_data: i32,
link: Node
}
and struct Node(Rc<NodeStruct>);
or struct Node(Option<Rc<NodeStruct>>);
if I need cross-links.Great thing is you can then add 'methods' to your type with
impl Node {}
Or define operators and other traits with: impl Add<Node> for Node {}
Sometimes, when I need mutability I even wrap the NodeStruct in RefCell.It seems like a lot of wrappers but thanks to them you can have very nice code that uses this type that has pretty much 'normal modren language' semantics + value semantics and is still blazing fast.
When you implement Ord, Eq, Hash they all go through all the wrappers and let you treat your final type Node as a comparable, sortable, hashable and cheaply clonable value. Dereferencing also goes through all or most of the wrappers automatically.
Main objects in my program are expression trees. I manipulate them, cut them, merge them, compare them, splice one into the other. Rc's enable me to have full flexibility and share tremendous amount of data across objects in my program.
Rust is absolutely wonderful language for this problem thanks to Rc's, enums, value semantics, auto-deriving traits and ability to implement traits for existing types and of course speed.
I'm not implementing specific algorithms. I'm making them up as I go although I used some simple ones like topological sort or A* that eventually turned into just breadth search because I have no idea how far I am from the solution.
It's mindboggling to me that people are using a systems programming language for mathematical research, especially if they don't know yet what the final algorithms will look like.
But all the more power to you for trying.
Some complain about this, but the fact is there's no such thing as a zero-overhead "copy" for non-trivial types. C++ started out with = meaning clone the object which was an even bigger footgun, and support for move had to be added after the fact.
It's very elegant solution for simple, small data types. But it further occludes how meaningfully Rust is different from everything else because thanks to that = sometimes does mean copy.
You need to structure your program as a Directed Acyclic Graph (DAG), with things interacting only with the things below them in the graph.
Then occasionally you might need to break the DAG structure by using Rc, Cell & RefCell, etc...
And finding it towards the end of writing your program after hours of fighting with borrow checker is extremely unpleasant.
And I don't think I ever landed in the situation where I could fix the discrepancy by sprinkling in few Rc, RefCells and such.
So I prefer to write with RefCells from the start and when I got the thing working and I am ambitious enough then I look at which parts could be borrows instead and I swap them out.
The issue here is that you are writing C++ code rather than Rust code.
Two separate synced trees? Is it worth it?
> The issue here is that you are writing C++ code rather than Rust code.
How dare you! I'm writing TypeScript code! ;-)
Rust is not Forth. I can write whatever I want and there's nothing wrong with that.
Rewrite your program in a form where it does not contain a tree.
If you want an actual tree as a data structure, see the trees crate.
> Rust is not Forth. I can write whatever I want and there's nothing wrong with that.
And other people write Haskell code in Python :p. If your code style doesn't match the language you are using you are going to have a lot of unnecessary friction.
But you inspired me about something. I think I can rewrite the program that I am writing to use reverse Polish notation instead of a tree. Thanks!
Though that should not apply to Rust at all, as it does not pledge to follow the C ABI internally (aka `extern "Rust"`).