The real gain is with MPI, anyway... shared memory is a trap. All the projects I've worked on have exploited parallelism through message passing, or a battle tested multithreaded linear algebra library.
What? Lots of parallelism doesn't use message passing!
Hence, in order for that up-front cost to be worth it, the algorithms run on the GPU need to re-use that data many times. If the algorithms running on the GPU don't make extensive reuse of the data sent to it, it would be faster to just do the calculation on the host CPU.