This is precisely what vector::emplace() solves, and std::move should be faster than swap and pop. Modern C++ has changed a lot, this article ignores the massive improvements added in c++11,14,17.
This is precisely what vector::emplace() solves, and std::move should be faster than swap and pop. Modern C++ has changed a lot, this article ignores the massive improvements added in c++11,14,17.
The whole swap-and-pop section weirded me out. Maybe I just don't know enough about C++, but saying that assignment (a[i] = a[n-1]) will call the destructor seems false.
As far as I know, the compiler should generate an implicitly defined copy assignment operator for these fixed size PODs and it should be as performant as memcpy.
But again, I don't have years and years of in-depth C++ experience, so I would be grateful if an expert could shed more light on this.
I think that would just call the copy assignment operator, would it not?
For correctness you would probably then follow up with a pop_back to keep the vector right-sized.
Actually you'd probably want to do:
a[i] = std::move(a[n - 1]);
Then follow up with pop_back.
Best would probably be:
a[i] = std::move(a.erase(n-1));
a[i] = std::move(a.erase(n-1));
There is no erase that takes an index, so I assume that n = a.end(). Also it is missing a dereference: a[i] = std::move(*a.erase(a.end()-1));
but erasing the one-before-the-end returns the (new) end iterator, which obviously is not referenceable. In general, after calling erase, it is too late to access the erased element.You want something like:
template<class Container, class Iter>
auto erase_and_return(Container&& c, Iter pos)
{
auto x = std::move(*pos);
c.erase(pos);
return x;
}
Also in the general case it doesn't make sense for erase to return a move iterator.As per the author's constraints these are "POD types that are trivially memcpy-copyable", so by definition the copy constructors will never do anything. Much less "allocate memory" as the author claims.
I'm not a game developer, but have spent a decade doing C++ on Windows, and at former employer, we had several different debugging profiles depending on the severity/difficulty of reproducing/debugging an issue. Our "normal" debug profile had all of the debug checks in the std lib disabled, and we could only effectively debug our own code. Not sure if games dont do this, or if its still not performing enough.
At work we don't use a debug build in the traditional sense, it's what you call a no-optimisations build where the code is compiled without most optimisations but otherwise the flags are the same as a release build. Some teams also go a step further and compile most of the code in release but some of their code with optimisations disabled.
They don't have to be, but it certainly makes this world's easier. If the flags are not the same, for sure you have to be very careful about passing objects between DLL boundaries.
At the companies I've done C++ work at, we've always had the source for all non C libs and compiled any C++ libs our selves (except for Windows libs, bit they also provide checked debug libs), so we could control the flags.
- Debug build performance. Release builds of C++ code using STL are generally pretty fast, but Debug builds suffer a lot (especially Visual Studio's std::vector implementation is notoriously horrible for debug builds). Debug executable speeds matter when you are debugging a game; you don't want to test your first-person shooter in 1 FPS!
- Build speed. Because of heavy use of templates and historical cruft, STL slows down your build times a lot. The build-test cycle is very important when designing games; you don't want to wait for a few hours after you've changed a few lines of code to tweak a new feature. Gigantic distributed build servers alleviates this problem a bit, but they are pretty cumbersome to set up nonetheless.
Though if you'll be calling the same function repeatedly to accumulate content into a single container it is far more efficient to have a function with an output reference rather than returning a new container. This will result in fewer memory allocations and you can also pre-allocate the size once before calling those functions.
On the part of tooling it might be nice if there was a way to annotate a function so that it creates a warning if the compiler cannot use copy-elision for the return value. (To be honest I haven't checked the documentation for this specific thing)
edit: also, the rule is simple: RVO is always mandated, NRVO remains an optimization.
Using POD structs which can be zero-initialized and memcpy'ed may indeed be faster, especially when these are bulk-operations.
There are plans to solve this - http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p114...
std::vector itself falls in this category: trivially relocatable, definitely not trivially copyable. So a vector of vectors will not necessarily be able to use memcpy but rather fall back to copy/move assignment. This is not very significant in performance for this type (vector move being cheap) but a language gotcha nonetheless (as the move constructor will be called n times in every capacity change)
What unfortunately is not defined is (trivially) relocatable as that's not a property that can be safely be inferred so it is not (yet) part of the standard. Some libraries still have this concept and require some sort of opt in.
If you're not careful, it will call an implicit constructor.
> So in general, if both push_back() and emplace_back() would work with the same arguments, you should prefer push_back(), and likewise for insert() vs. emplace().
The reference returning emplace_back() is used frequently in the code to construct a new element of a struct and then fill in its members, as opposed to creating a new struct then push_back() to copy the memory in.
Similarly, std::vector has to allocate more memory every time it has to grow and copy all its contents, whereas for POD datatypes you can just use realloc which can save copies.
These are all borderline microoptimisations, but they matter for realtime highly responsive software. Or just in general when you need to squeeze out every last bit of performance.
If you want an `A` struct/class, you'll call the ctor/dtor, that's true. But for POD types, if the ctor/dtor does nothing, they are trivial to inline and will incur no runtime overhead by any compiler nowadays.