An evaluation of throughput computing on CPU and GPU (2010) [pdf]
sbel.wisc.edu
sbel.wisc.edu
The thing is, this is still a real win for CUDA, because most people aren't going to write highly optimized C. It's just not a hardware win, it's a compiler win. The thing that most people get out of CUDA is not an advantage from running on a GPU, but an advantage of a smart compiler/language design. Most of the same people would get all but 1-5x of the benefits from using something like ISPC[1], which gives a similar programming model as CUDA implemented on the CPU.
In my experience (from the Tesla/GTX 200 era, though I doubt things have changed that much now), such a small boost in performance is not worth the hassle of transferring data to/from the GPU, the (lack of) virtualization, and driver shenanigans/support issues (at one point, I had to suggest someone buy a fake monitor dongle to plug into his GPU to be able to use my code...).
Anyways, it's good to see this thread, people these days think you are crazy for not using the GPU to do big compute workloads.
There are a few more quirks to keep in mind and you need to have the memory lay-out and access patterns down otherwise it will not give you the boost you expect.
In a way I really love CUDA for this reason, it allows me to re-cycle all my old optimization skills and do something useful with them.
Highly optimized C and CUDA are, to me, the same concept - to get the best results, you have to understand the underlying machine to get the best results. While your resources are different, it's a very similar skill - to codesign an algorithm with the deep knowledge of how either a modern CPU or a GPGPU actually works. I'm no longer current as a GPGPU programmer but it wouldn't surprise me that the same 10-20x that an expert can get with astute design for modern architecture is available on GPGPU as well (i.e. the difference between a naive and a sophisticated design on the same platform).
Personally I regard that kind of algorithmic codesign as one of the most fun things you can do with your clothes on, but that's probably just me.
Some problems are well suited to this kind of architecture and some not, if you're in luck the gains can be significant and it is definitely something you should be aware of and weigh when you're writing something that is heavy on computation.
The CPU implementations had had significant optimisation work done on them because it had to run on 10k+ cores so even small improvements had significant cost implications.
The GPU implementation ran about 40-50x faster for the same accuracy. We didn't need that speed up in practice so we traded some of it for more paths.
"We show that CPUs and GPUs are much closer in performance (2.5X) than the previously reported orders of magnitude difference."