GPU is the way to go these days to get most parallelism, and it is definitely the shared memory variety. Message passing is the go to when parallelism needs to cross machine boundaries.
Hence, in order for that up-front cost to be worth it, the algorithms run on the GPU need to re-use that data many times. If the algorithms running on the GPU don't make extensive reuse of the data sent to it, it would be faster to just do the calculation on the host CPU.