Intrinsics are directly callable from C++.
> likely/unlikely branches
Most compilers have extensions that will allow you to do this (__builtin_expect and so on).
> in-lining can't be forced when you know it gives better performance
Again, most compilers have this, not just GCC, e.g. __forceinline.
> the compiler has a lot of trouble knowing when lines of code are independent and can be done in parallel (b/c const =/= immutable)
This is true, as aliasing is a real issue. The hardware itself has some say over this anyway, dependent on its instruction scheduling and OOE capabilities.
What you don't mention, however, is the fact that almost no other languages offer any of these, let alone all of them. Rust may be the exception here, although some of this is still in the words (SIMD, I'm not sure about the status of likely/unlikely intrinsics).
For GPU programming, if you're using CUDA, you're almost certainly using C or C++, or calling something that wraps C/C++ code. Not everything is suited to GPU processing anyway, there's still a lot of code that's not moving off the CPU any time soon that needs to be performant.