Throughput 100% depends on how hard they need to hit memory. GPUs love stuff that they can do entirely within registers. Having to hit (the limited quantity of) shared memory slows things down much more. Inefficient/non-coaslesced use of shared memory or global memory too often can trash performance. Hitting host/CPU memory in any non-trivial amount usually dooms a program.
With each step you not just cut down your bandwidth but also you increase your latency and that has a huge impact.