Short intro to C++ for Rust developers: Ownership and Borrowing
nercury.github.io
nercury.github.io
Returning a stack allocated value does not invoke the move constructor. It doesn't copy either. This is a separate optimization that existed even prior to C++11 called the "named return value optimization" or NRVO for short. And by "sufficient new compiler" you really should just say every compiler. No compiler I know of in existence used by actual people doesn't implement this.
Here's a demo: http://ideone.com/QRqj5P
It's VERY important that this optimization isn't confused with move construction as the latter would actually be extremely less performant than what actually happens.
Aren't copy elision guarantees part of the new C++17 standard?
Returning a value from a C++ function is a "move" in the Rust sense. It's not a move in the C++ sense, since C++ moves involve move constructors.
It's probably this confusion that led to it being written that way, probably should be fixed.
return val; // val is `rvalue` here
This is the part that isn't accurate. "val" is a named reference and is not an rvalue. There isn't any "moving" up the stack that happens. He's likely confusing this with an expiring value (a subtype of rvalue) that happens in code like foo(returns_something());
The returned value from returns_something() is un-named and is about to die after this line executes (lives until foo's invocation unwinds) and so it is an rvalue. You can't have a named rvalue unless you std::move explicitly I don't think.This is the one time in the language that lvalues are automatically moved from, and it's because the language/compiler knows that returning from a function will destroy all local variables, so it's safe to move from them.
Giving the compiler the ability to move implicitly surprises me since moves can have side-effects as you've said.
A lot of C++ resources explain moving solely in terms of rvalues so it's easy to associate the two in the wrong places.
NRVO is a compiler optimization, not a feature of the language itself.
So it's up to the compiler to elide the copy (making it a compiler optimisation as you say), but the standard says that it has to move if it doesn't elide. This means that returning a named loca variable by value is always the best choice. Again, see "Effective Modern C++" for more.
trait Copy: Clone {}
Copy is nearly always done by memcpy/memmove. When Copy is invoked it won't use the `Clone::clone` method: https://is.gd/2hMyNP.
i.e. Clone is always explicit.Added: Essentially Copy is the same as a move, but the original value is still valid.
Rust doesn't, so it's just an optimization. In C++ the code will have different semantics (e.g. if your move ctors have side effects) based on this optimization, so the language needs to define if/how it works.
I have removed a bit where I say that rvalues are same as returned values, to avoid confusing other readers.
However, if the name was a very big string, i.e. something like file contents, and it would be necessary to ensure no copies for performance reasons, the unique_ptr or shared_ptr would come to the rescue
This is also not true at all. You can choose to provide a non-templatized function that only accepts an rvalue reference. For example, http://ideone.com/055ViY does not compile.The purpose of the smart pointers unique_ptr and shared_ptr is to annotate ownership for heap allocated objects (unique vs shared ownership). They do help alleviate the thing you're describing, but shouldn't be used so liberally in that way since they heap allocate!
See also: http://en.cppreference.com/w/cpp/language/copy_elision
Unless the 'named return value' you're returning happens to be a named parameter.
Example: https://ideone.com/TddXn7
The move here actually comes from the return, not passing in the newly constructed object. The compiler is eliding Foo's move constructor in main(), but it can't elide both... even if the function is fully inlined.
In your case a move would happen but if you changed the parameter to a reference I think it would copy.
This results in less code and more manageable semantics.
Consider this example: http://ideone.com/T872Oj
In F1 the compiler can NRVO away the move.
In F2 it cannot.
Both will fail to compile if you delete the move constructor.
Edit: I see that StephanTLavavej clarified that below.
I'm not saying that you are wrong, only that your demo poorly demonstrates what you are arguing.
This article actually helped me understand C/C++ a little bit better than I did before. For example where it explained the differences between Rust and C/C++ copy vs move behavior... I got a better sense of just how C/C++ works and why I might chose to write, say, a function parameter one way vs. another.
So in this basic sense, mission accomplished I think. I have no doubt there's lots of nuance missing, but still... not bad and I appreciate the effort.
In most cases maintainability wins,
and avoiding “premature optimization”
is very much a necessity in C++.
I agree with this conclusion. Quite often, when you start coding you don't know, where the performance bottlenecks hide, and you don't want to waste time, thinking about allocating memory for a routine which you could write equally well in any scripting language.
Unfortunately plugable automatic garbage collection for less verbose performance uncritical scripting tasks within the language is not usable.Is
void foo(T&& value) {
...
}
similar to fn foo(value: T) {
...
}
?Or at least that's the idea. In C++, it's up to the author of T to implement such semantics (through its move constructor) correctly for non-trivial types.
Modern C++ is pretty darn nice to work with. It's not a "safe" language by any means, but with good tooling, design, and a good static analyzer you can prevent the vast majority of memory bugs you'd run into with C. Plus the compilers are so mature you get great performance right out of the box.
Yup. And if you don't you can drop all the way down to intrinsics and assembly. It's pretty nice.
Actually Windows has always enjoyed better C++ related tooling than any UNIX other than Mac OS X.
And standards compliance is actually the best one among commercial C++ vendors.
http://en.cppreference.com/w/cpp/compiler_support
Also there is the world of commercial UNIXes, mainframes and embedded OSes, where using clang or gcc isn't always an option.
For example TI is still stuck on C++98, not even C++03.
Also, there's no qualitative difference between C++ toolchains on macOS and Linux, since the same compilers are used on both.
> no qualitative difference between C++ toolchains on macOS and Linux,
Where is Instruments, XCode, Cocoa, IO Kit for GNU/Linux?
- graphical debugging of parallel tasks and threads
- displaying data structures graphically with code navigation
- WYSIWYG UIs with tooling like Blend
- GPU debugging
- Code navigation across object files, shared libraries and source code
- Incremental compilation and linking with intermediate representations stored in a database.
- Out of the box representation for all STL data structures and user defined types
Tooling is much more than just a plain old compiler.
What Microsoft compiler doesn't provide in my experience is decent codegen.
Not when compared with the enterprise and ultimate versions. Or with the changes done in VS 2015, improved in the upcoming VS 2017 for incremental building and database storage of symbols.
> Both QT and GTK have GUI builders.
Miles behind of what Blend + XAML allow for.
> NVidia provides a plugin to debug and profile GPU code.
How well does it work with Intel and AMD cards?
> What Microsoft compiler doesn't provide in my experience is decent codegen.
On Windows, among commercial compilers, only ICC generates better code.
There are reasons for bad code navigation and missing refactoring features: https://blogs.msdn.microsoft.com/vcblog/2015/09/25/rejuvenat...
Ha I always thought there were actually none.
even find all references is not reliable
Yup that sucks. It's better in VS2015
Project files are a mess.
Yeah it's xml but for the rest it just lists files and options, how is that much of a problem?
There is vcxproj file to define build and vcxproj.filters file for directory structure view in VS (I do not understand a reason for this).
Can be easily solved by ditching the filters file all together and enabling 'All Files' option in Solution Explorer. Unless you insist on having all files listed by extension, which is imo a useless complete mess for enything but small projects.
you can't set hexadecimal format just for one variable
You can, use '<variablename>, h' in the watch window
Does not work for me, because files are located in subdirectories. I see only long list of files from all directories without filters. Maybe it is just bad project structure but I cannot change it. I found plugin for 2015 which is able to generate filters from directory structure, but I have to wait for upgrade.
Not if you select 'Show All Files' in solution explorer, then it shows the entire subdirectory tree.
Depends on what you mean with tooling I guess. CLang on windows is fine these days, so is msys. So, with possibly some quirks, that unix tooling runs on windows as well so it cannot be vastly worse tooling if it's the same.
Still, talking IDEs and visual debugging which you might or might not consider a part of tooling, unfortunately VS still runs circles around the free competition. So much that I'm betting teams are willing to give up some standard compliance in favour of using it.
It's a major pita when MSVC cannot understand a pefectly standard C++ construction.
And Microsoft is well aware of that, they made jumps and leaps recently to bring their level of C++ compiler support to the actual standard level, finally you can say MSVC actually understands standard C++.
Never use "std::move"
It's really just there to allow fancy optimisations and you don't need it in most cases. You should rely on the sane defaults instead:
- Putting variables on the stack to manage object ownership/lifetime is a good default.
- Use std::unique_ptr or std::shared_ptr() for heap ownership (single ownership vs. shared)
- For a non-owning (mutable) reference, use const & (&) or const * (*) when the value may not exist.
- Always prefer passing by value (return types and parameters)
- If a parameter type is heavy, take a (const) reference.
- If the return type is heavy, the compiler's RVO will do its job.
No need to be more fancy than this
Essentially you use std::move when you want to have pointer semantics without using a pointer. You don't want to use a pointers sometimes to keep objects on stack and use it to track ownership.
So instead of doing C* c1 = new C; C* c2 = c1; You have C c1; C c2 = c1; // here I have a move constructor which moves c1 to c2, c2 is now a null object.
This is what rust is able to do and you don't need expensive atomics (shared_ptr), ugly template syntax (unique_ptr), can keep everything on the stack and (!) have the objects be automatically cleaned up correctly.
You should definitely understand it, but if you find yourself wanting to use it, you should start by considering alternatives first. It's a sharp (and often unsafe) tool that should be avoided in most cases.
For your example, unique_ptr is the correct (and safe) solution and will do the same thing.
> you don't need expensive atomics (shared_ptr)
Note that shared_ptr provides no "atomicity". It is in reality exceptionally cheap.
Person(std::string first_name, std::string last_name)
: first_name(std::move(first_name))
, last_name(std::move(last_name))
{}
For string parameters, do you need to use std::move()? Won't the compiler do that anyway?And does anybody really put the comma separating initializers at the start of the line? Yuck.
Hangover from SQL as well.
If the language supports trailing commas[0] (rather than just interspersed ones) that's not an issue as you could write:
Person(std::string first_name, std::string last_name):
first_name(std::move(first_name)),
last_name(std::move(last_name)),
{}
[0] as Python, Ruby, Javascript or Rust doI don't, but I can see why someone would. When you copy-paste initializer lines around (say, because you've re-ordered members in the declaration), a leftover trailing comma often leads to insanely painful compiler errors. Organizing the commas this way makes that less likely.
For example, suppose you're writing a FancyAppend() function, that will return "LhsString, RhsString". Following your guidance, the signature would be "string FancyAppend(const string&, const string&)". While that definitely avoids copying the inputs, it is not optimal. Providing additional overloads for string&& parameters can be more efficient (if at least one input is a modifiable rvalue, it can be appended to in-place, which can avoid additional memory allocation if it happens to have sufficient capacity; for repeated appends, this can avoid quadratic copying of elements). Indeed, this is what string's operator+() does.
Note that working with rvalue references does require more understanding. For example, the signature "string&& BadAppend(string&&, const string&)" is a severe error (one that the Standardization Committee made early on, and corrected before shipping).
But C++ is like a large buffet where you can pick from a lot of different abstractions.
Rust on the other hand is not as flexible. But it rather tries to focus in safer abstractions, and provide convenience and coherence around them.
How so? If you want the same flexibility as C++ memory safety wise, you just use the unsafe keyword when you need to. A lot of Rust programmers are scared of unsafe blocks for some reason but writing unsafe code in Rust is no less safe than writing regular C/C++ except you can choose to reenable safety whenever you want to. You have to look out for subtle bugs that come up when safe code makes assumptions about unsafe code that you don't properly implement but these are logic bugs, same as you'd find in any other language when a module breaks a contract. Just don't use unsafe willy nilly in a library crate that you expect other people to use.
As far as the rest of the language, I find Rust's traits to be far more flexible than C++ OOP, it just takes a little while to adapt to thinking in trait based composition. Once trait specialization and "impl Trait" hit stable and the Rust team figures out how to make the ambiguity rules less conservative, we'll have the best of both worlds: interfaces (both dynamic and static), trait based composition, and inheritance (you can already do regular OOP inheritance with a little bit of boilerplate like how Servo uses Upcast<T>/Downcast<T> traits for HTML object hierarchies). Both are currently in nightly (impl Trait might be beta already).
Many unsafe abstractions are not opt-in in C++. You can go ahead and writing them. No warning or cost.
"mut" and "unsafe" are keywords that you need to write down each time. They're keywords you can spot easily during a code review and say "how do you justify this?".
Rust's language design choice of having everything being immutable and safe by default is a great idea.
Yes, but that is scoped. You can write unsafe code when designing an abstraction, and seal off the unsafety with a safe-to-use API. You then verify the abstraction, and are free to use it however you want. This is what's usually done, and this is what unsafe is for.
"mut" isn't really a red-flag keyword. Sure, you don't want unnecessary mutation and everything is immutable by default, but nobody really has issues with making things mutable. Rust is not a language where purity is important.
"unsafe" is, but for designing abstractions that's what unsafe is for, and while you still should justify it, "I need an abstraction like this and it can't be implemented in safe Rust because <reasons>" should be enough. Of course, you should be able to explain why it is safe to use, preferably in the docs/comments.
So these abstractions may be implemented unsafely, but once implemented you don't need unsafe to use them.
The only abstractions that C++ has that Rust doesn't are those involving intrusive datastructures (and you can make them work in some cases).
(That by employing pointers or sentinels.)