I could be totally wrong, but have you had to optimize a tight loop in C/C++? It sucks and a lot of stuff is missing : simd, likely/unlikely branches, the compiler has a lot of trouble knowing when lines of code are independent and can be done in parallel (b/c const =/= immutable), in-lining can't be forced when you know it gives better performance (if you use GCC extensions you can alleviate a lot of this..but that's tying you to a compiler). Yeah C++ is generally faster than Go, but there is a lot of performance typically left on the table. The language kinda start to work against you when you start to dig down. So you slap some shit together and the compiler will do a decent job, but most C++ devs aren't even looking at the compiler output
But the elephant in the room is that more and more of our available flops are on the GPU and C++ isn't helping you there at all. Not only that, but the GPU is giving you way more operations per Watt (and that's what a lot of those people care about). And finally, when you throw stuff onto the GPU you are also leaving the CPU available to do other things. So there are a lot of "wins" there. As you illustrate, the areas of C++'s relevance is shrinking, and shrinking into the area that is very GPU friendly.
So the way I see it, C++ folks will start to write more OpenCL kernels for the performance critical pieces and the rest won't matter (Go or Clojure or whatever). The GLSL is kinda lame and too C99ish... so maybe someone will write a better lang that compiles to SPIR-V, and it's not exactly write-once-run-everywhere, but it could be much better than writing optimized C++ and it can run everywhere. It's more of the cross-platform-assembly C/C++ wants to be