- How do you deal with software that has been previously run with coarse grained parallelism, optimized for Multicore/Multinode x86? In my experience, GPGPU porting often leads to a tedious, mostly mechanical conversion from coarse grained to fine grained, which usually includes privatizing all your data manually in your parallel domains.
- Can you do multinode / multi-GPU without wrapping everything in MPI?
- Can ArrayFire also run on CPU clusters?
- How do you deal with different storage orders? Up until now, GPUs often require a different storage order than CPUs, (wide vs. narrow vector processor) - and how does that factor into the last point?
Sidenote: I've been dealing with above problems in a Fortran based research project and have created a preprocessor framework[1] to deal with it.