Disclaimer: I am not very experienced in GPGPU field, so my worrying may be proved wrong.
Disclaimer: I am not very experienced in GPGPU field, so my worrying may be proved wrong.
but what you are perhaps missing is that it's ok for gpus to read memory, as long as you have enough threads. they can switch context very quickly, so one set of threads can request memory (hopefully a contiguous chunk) and then drop into the background and let another set of threads do some work (on the same processing unit). this is critical to their efficiency and is very different to a cpu, which instead relies on cache and "sits doing nothing" if it needs to read data from "afar" (obviously there are trade-offs - there's only so much local memory for state, for example).
i worked on a problem that was not as "nice" as you might hope - the memory access was unpredictable to some degree. but i still got a speed up of "tens" on a cheap ($200) graphics card, compared to a meaty xeon. it's more robust than you might expect.