There is transfer time. You source image is in cacheable CPU memory. Integrated GPUs normally work with uncacheable memory allocated in a special region of system memory. Some GPUs can access cacheable memory too, but it is much slower (because it has to maintain coherency with CPU caches), and requires that you allocate such cacheable memory using special OpenCL driver calls, not your normal malloc. So, in practice, you would do a copy to GPU-optimized buffer (in shared with CPU, but uncacheable memory).