GPU parallelism is highly specialized and requires very specific workloads (e.g., many linear algebra workloads). GPUs are a bad choice for general parallelized tasks.
Hence, in order for that up-front cost to be worth it, the algorithms run on the GPU need to re-use that data many times. If the algorithms running on the GPU don't make extensive reuse of the data sent to it, it would be faster to just do the calculation on the host CPU.