- Debug build performance. Release builds of C++ code using STL are generally pretty fast, but Debug builds suffer a lot (especially Visual Studio's std::vector implementation is notoriously horrible for debug builds). Debug executable speeds matter when you are debugging a game; you don't want to test your first-person shooter in 1 FPS!
- Build speed. Because of heavy use of templates and historical cruft, STL slows down your build times a lot. The build-test cycle is very important when designing games; you don't want to wait for a few hours after you've changed a few lines of code to tweak a new feature. Gigantic distributed build servers alleviates this problem a bit, but they are pretty cumbersome to set up nonetheless.
I'm not a game developer, but have spent a decade doing C++ on Windows, and at former employer, we had several different debugging profiles depending on the severity/difficulty of reproducing/debugging an issue. Our "normal" debug profile had all of the debug checks in the std lib disabled, and we could only effectively debug our own code. Not sure if games dont do this, or if its still not performing enough.
At work we don't use a debug build in the traditional sense, it's what you call a no-optimisations build where the code is compiled without most optimisations but otherwise the flags are the same as a release build. Some teams also go a step further and compile most of the code in release but some of their code with optimisations disabled.
They don't have to be, but it certainly makes this world's easier. If the flags are not the same, for sure you have to be very careful about passing objects between DLL boundaries.
At the companies I've done C++ work at, we've always had the source for all non C libs and compiled any C++ libs our selves (except for Windows libs, bit they also provide checked debug libs), so we could control the flags.
Though if you'll be calling the same function repeatedly to accumulate content into a single container it is far more efficient to have a function with an output reference rather than returning a new container. This will result in fewer memory allocations and you can also pre-allocate the size once before calling those functions.
On the part of tooling it might be nice if there was a way to annotate a function so that it creates a warning if the compiler cannot use copy-elision for the return value. (To be honest I haven't checked the documentation for this specific thing)
edit: also, the rule is simple: RVO is always mandated, NRVO remains an optimization.