The thing is, this is still a real win for CUDA, because most people aren't going to write highly optimized C. It's just not a hardware win, it's a compiler win. The thing that most people get out of CUDA is not an advantage from running on a GPU, but an advantage of a smart compiler/language design. Most of the same people would get all but 1-5x of the benefits from using something like ISPC[1], which gives a similar programming model as CUDA implemented on the CPU.
In my experience (from the Tesla/GTX 200 era, though I doubt things have changed that much now), such a small boost in performance is not worth the hassle of transferring data to/from the GPU, the (lack of) virtualization, and driver shenanigans/support issues (at one point, I had to suggest someone buy a fake monitor dongle to plug into his GPU to be able to use my code...).
Anyways, it's good to see this thread, people these days think you are crazy for not using the GPU to do big compute workloads.